SkillGym-Qwen3.5-9B

SkillGym-Qwen3.5-9B is Qwen3.5-9B finetuned on SkillGym trajectories to use agent skills. It is one of the models released with the paper SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation.

An agent skill is a folder of instructions, reference documents, and scripts that an agent can consult while solving a task. This model is trained to read the relevant skill and apply it, through tool calls or shell commands, in a sandboxed workspace.

Results

Scores are on a 0-100 scale, higher is better. SkillGym is task success on the SkillGym test set. SkillEval and SkillsBench report the mean and standard deviation over three runs. Skill-Use-Bench reports the skill-use (SU) score. All models are evaluated with skills available, using MiniSwe-Agent as the harness.

Model Size SkillGym SkillEval SkillsBench Skill-Use-Bench
MiniCPM5 2B 33.2 63.0 ± 0.6 10.8 ± 2.7 24.1
+ SkillGym SFT 2B 33.2 68.1 ± 2.0 6.1 ± 0.8 44.7
Ministral-3 8B 17.5 55.7 ± 3.1 3.9 ± 2.7 4.3
+ SkillGym SFT 8B 49.5 77.5 ± 0.4 16.9 ± 1.3 55.1
Qwen3.5 4B 33.8 62.8 ± 1.5 10.1 ± 1.3 8.8
+ SkillGym SFT 4B 47.0 70.8 ± 1.6 14.3 ± 1.0 48.7
Qwen3.5 9B 41.3 65.7 ± 1.1 14.8 ± 3.5 14.8
+ SkillGym SFT (this model) 9B 59.5 74.7 ± 1.2 22.4 ± 4.2 49.6
Qwen3.5 27B 55.3 76.2 ± 1.1 32.6 ± 2.3 28.3
+ SkillGym SFT 27B 62.8 80.1 ± 1.5 47.4 ± 5.1 74.2
Qwen3.5 122B/10B 53.0 69.8 ± 1.1 30.1 ± 2.5 16.8
+ SkillGym SFT 122B/10B 65.0 80.0 ± 0.3 53.6 ± 5.2 72.2

Training

Base model Qwen/Qwen3.5-9B
Data reasonwang/skillgym-sft, 19k verified successful trajectories from three teacher models
Method Supervised finetuning, 2 epochs
Learning rate 1e-5 with linear decay, AdamW
Batch size 128
Max sequence length 65,536 tokens
Reasoning Trained with reasoning traces

Usage

The model keeps the chat template of its base model and can be served with any OpenAI-compatible server, for example vLLM.

vllm serve reasonwang/SkillGym-Qwen3.5-9B

In our evaluation we use temperature 0.6, top_p 0.95, top_k 20. The skill names and descriptions are listed in the system prompt, and the agent reads the skill files itself.

Limitations

  • The model is an agent policy for tasks with skills and tools. It is not tuned as a general chat assistant.
  • Long episodes can exhaust the context budget before the task is finished.
  • Reasoning occasionally repeats itself and runs to the per-turn token limit.

Citation

@article{wang2026skillgym,
  title   = {{SkillGym}: Training Skill-Use Agents with Automatic Verifiable Environment Generation},
  author  = {Wang, Renxi and Hee, Mingshan and Koto, Fajri and Baldwin, Timothy and Li, Haonan},
  journal = {arXiv preprint arXiv:2609.37539},
  year    = {2026}
}
Downloads last month
135
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for reasonwang/SkillGym-Qwen3.5-9B

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(989)
this model
Quantizations
3 models

Dataset used to train reasonwang/SkillGym-Qwen3.5-9B

Collection including reasonwang/SkillGym-Qwen3.5-9B

Paper for reasonwang/SkillGym-Qwen3.5-9B