qwen3-8B-length-penalty-global-ref-iter259
GRPO-trained deep-research agent on BrowseComp-Plus (search / open_page /
finish tools over a fixed 100K-doc corpus). Cell qwen3-8B/length_penalty_global_ref of the
length-penalty experiment matrix, checkpoint at training iteration 259.
- Code: https://github.com/ys-2020/miles (branch
browsecomp-rl, seedocs/experiments/browsecomp-length-penalty-results.mdfor the full study) - WandB: project
browsecomp-b300 - Trained with miles fully-async GRPO: group size 8, global batch 256, temp 1.0, lr 1e-6, KL 0.001, ~100-turn ReAct rollouts, 40960-token context.
Offline eval (150 BrowseComp-Plus test questions, temp 0.6, 1 sample/q)
| iter | accuracy | mean response len (tokens) | truncated ratio |
|---|---|---|---|
| 19 | 0.247 | 2334 | 0.22 |
| 39 | 0.240 | 2244 | 0.20 |
| 59 | 0.260 | 2325 | 0.21 |
| 79 | 0.280 | 2285 | 0.13 |
| 99 | 0.300 | 2479 | 0.17 |
| 119 | 0.273 | 2541 | 0.17 |
| 139 | 0.333 | 2280 | 0.14 |
| 159 | 0.340 | 2339 | 0.07 |
| 179 | 0.373 | 2143 | 0.06 |
| 199 | 0.340 | 2026 | 0.05 |
| 219 | 0.360 | 1986 | 0.02 |
| 239 | 0.407 | 1988 | 0.01 |
Resuming training in miles
Convert back with tools/convert_hf_to_torch_dist.py, rename the saved dir to
iter_0000259, write 259 into latest_checkpointed_iteration.txt, then
launch with RESUME=1 (see examples/browsecomp/slurm_gb300/).
training_metadata.json in this repo records provenance.
- Downloads last month
- 20
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support