Qwen3-8B, GRPO-trained on search_plan_act
Full-weight checkpoint (LoRA merged) of Qwen/Qwen3-8B, trained with GRPO
(via TRL's multi-turn environment_factory tool-calling API) on
search_plan_act: a procedurally-generated, closed-book multi-turn task
requiring the model to search for facts, read_record/update_record/
link_records on retrieved entities, and finish, looping through tool
calls until the goal is achieved or it gives up. Reward decomposes into
outcome / grounding / stop-behavior components, with an explicit
adversarial-pilot check that no reward-hacking strategy (reckless guessing,
spamming searches, never finishing) scores above genuine goal-directed play.
Training: LoRA (r=32, alpha=64) on all attention+MLP projections, GRPO via TRL 1.9.2 + vLLM colocate generation, 250 steps, effective batch size 32, 4x GH200 GPUs.
Evaluation
Evaluated pre-RL (base Qwen/Qwen3-8B) vs post-RL (this checkpoint) on held-out
real-world and synthetic benchmarks, all using vLLM-accelerated generation
with the model's own chat template (apples-to-apples generation path in both
rows):
| Benchmark | Pre-RL | Post-RL (this checkpoint) |
|---|---|---|
| MuSiQue (EM / F1, n=2417) | 0.108 / 0.191 | 0.338 / 0.456 |
| 2WikiMultihopQA (EM / F1, n=12576) | 0.585 / 0.681 | 0.585 / 0.681 |
| search_plan_ood (mean reward, n=100) | 0.0015 | 0.070 |
| BALROG BabyAI (mean episode return, n=50) | 0.000 | 0.020 |
MuSiQue and search_plan_ood/BabyAI show clear, real transfer from RL training; 2WikiMultihopQA is essentially unchanged.
Files
model.safetensors,config.json,tokenizer.json,tokenizer_config.json,generation_config.json,chat_template.jinja-- the full merged checkpoint.adapter/-- the raw LoRA adapter (pre-merge), kept for reference.
- Downloads last month
- 81