Cedar Math: GRPO Models & Data Collection Final GRPO models (1.5B, 3B, 7B), their shared 17,005-row DAPO math training set, provenance and benchmark scores. • 4 items • Updated 10 minutes ago
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization Paper • 2608.17067 • Published about 1 month ago • 22
zbeeb/deepseek-r1-distill-qwen-14b-fast-math-r1-sft-10ep Text Generation • 841k • Updated May 28 • 16
zbeeb/deepseek-r1-distill-qwen-14b-fast-math-r1-sft-10ep Text Generation • 841k • Updated May 28 • 16