qwen3-8B-length-penalty-global-ref-iter259

GRPO-trained deep-research agent on BrowseComp-Plus (search / open_page / finish tools over a fixed 100K-doc corpus). Cell qwen3-8B/length_penalty_global_ref of the length-penalty experiment matrix, checkpoint at training iteration 259.

  • Code: https://github.com/ys-2020/miles (branch browsecomp-rl, see docs/experiments/browsecomp-length-penalty-results.md for the full study)
  • WandB: project browsecomp-b300
  • Trained with miles fully-async GRPO: group size 8, global batch 256, temp 1.0, lr 1e-6, KL 0.001, ~100-turn ReAct rollouts, 40960-token context.

Offline eval (150 BrowseComp-Plus test questions, temp 0.6, 1 sample/q)

iter accuracy mean response len (tokens) truncated ratio
19 0.247 2334 0.22
39 0.240 2244 0.20
59 0.260 2325 0.21
79 0.280 2285 0.13
99 0.300 2479 0.17
119 0.273 2541 0.17
139 0.333 2280 0.14
159 0.340 2339 0.07
179 0.373 2143 0.06
199 0.340 2026 0.05
219 0.360 1986 0.02
239 0.407 1988 0.01

Resuming training in miles

Convert back with tools/convert_hf_to_torch_dist.py, rename the saved dir to iter_0000259, write 259 into latest_checkpointed_iteration.txt, then launch with RESUME=1 (see examples/browsecomp/slurm_gb300/).

training_metadata.json in this repo records provenance.

Downloads last month
20
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shangy/browsecomp-qwen3-8B-length-penalty-global-ref-iter259

Finetuned
Qwen/Qwen3-8B
Finetuned
(1997)
this model