VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models Paper • 2609.04355 • Published 13 days ago • 9
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy Paper • 2609.28660 • Published 8 days ago • 13
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 12 days ago • 17
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents Paper • 2609.23377 • Published 11 days ago • 50
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 15 days ago • 30
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention Paper • 2609.21788 • Published 13 days ago • 13
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 13 days ago • 138
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 13 days ago • 35
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 17 days ago • 159
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 14 days ago • 57
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 14 days ago • 41
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 14 days ago • 44
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents Paper • 2609.17653 • Published 16 days ago • 45
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments Paper • 2609.19134 • Published 15 days ago • 101
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 16 days ago • 77
ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals Paper • 2609.16816 • Published 16 days ago • 11
MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control Paper • 2609.06251 • Published 26 days ago • 8
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 24 days ago • 376