Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 3 days ago • 277
Running 238 The ultimate guide to RL environments: building and scaling them in the LLM era 📝 238 Building and scaling RL environments for LLM training
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 16 days ago • 269
Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding Paper • 2608.25356 • Published 22 days ago • 20
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 24 days ago • 207
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published Aug 7 • 113
Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Paper • 2608.02831 • Published Aug 3 • 16
Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Paper • 2608.02831 • Published Aug 3 • 16
AudioRubrics Collection Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning: model and rubric dataset. • 2 items • Updated Jul 12 • 1
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published Aug 7 • 113
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published Aug 6 • 64
ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory Paper • 2509.04439 • Published Sep 4, 2025 • 2
TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning Paper • 2510.03519 • Published Oct 3, 2025 • 1
FlowBank: Query-Adaptive Agentic Workflows Optimization through Precompute-and-Reuse Paper • 2606.11290 • Published Jun 9 • 2