Qwen3-4B, GRPO-trained on search_plan_act

Full-weight checkpoint (LoRA merged) of Qwen/Qwen3-4B, trained with GRPO (via TRL's multi-turn environment_factory tool-calling API) on search_plan_act: a procedurally-generated, closed-book multi-turn task requiring the model to search for facts, read_record/update_record/ link_records on retrieved entities, and finish, looping through tool calls until the goal is achieved or it gives up. Reward decomposes into outcome / grounding / stop-behavior components, with an explicit adversarial-pilot check that no reward-hacking strategy (reckless guessing, spamming searches, never finishing) scores above genuine goal-directed play.

Training: LoRA (r=32, alpha=64) on all attention+MLP projections, GRPO via TRL 1.9.2 + vLLM colocate generation, 250 steps, effective batch size 32, 4x GH200 GPUs.

Evaluation

Evaluated pre-RL (base Qwen/Qwen3-4B) vs post-RL (this checkpoint) on held-out real-world and synthetic benchmarks, all using vLLM-accelerated generation with the model's own chat template (apples-to-apples generation path in both rows):

Benchmark Pre-RL Post-RL (this checkpoint)
MuSiQue (EM / F1, n=2417) 0.316 / 0.429 0.316 / 0.431
2WikiMultihopQA (EM / F1, n=12576) 0.552 / 0.656 0.550 / 0.655
search_plan_ood (mean reward, n=100) 0.000 0.000
BALROG BabyAI (mean episode return, n=50) 0.000 0.019

Honest finding: this 4B checkpoint shows only a small gain on BALROG and is essentially flat on every other benchmark, including the synthetic in-distribution-adjacent search_plan_ood task. Compare against the sibling bmonikraj/qwen3-8b-search-plan-act checkpoint, which showed much stronger, more consistent transfer under the identical training recipe -- RL transfer here is genuinely model-specific, not just a function of training recipe or model family.

Files

  • model.safetensors, config.json, tokenizer.json, tokenizer_config.json, generation_config.json, chat_template.jinja -- the full merged checkpoint.
  • adapter/ -- the raw LoRA adapter (pre-merge), kept for reference.
Downloads last month
3
Safetensors
Model size
4B params
Tensor type
BF16
·
Video Preview
loading

Model tree for bmonikraj/qwen3-4b-search-plan-act

Finetuned
Qwen/Qwen3-4B
Finetuned
(1052)
this model