Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures Paper • 2609.13463 • Published 14 days ago • 5
Studying Without a Syllabus: Task-Agnostic Environment Preprocessing Paper • 2609.10824 • Published 16 days ago • 4
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Paper • 2607.28802 • Published Jul 30 • 11
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions Paper • 2606.30573 • Published Jun 29 • 8
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR Paper • 2605.20164 • Published May 19 • 6
VDGD: Mitigating LVLM Hallucinations in Cognitive Prompts by Bridging the Visual Perception Gap Paper • 2405.15683 • Published May 24, 2024
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities Paper • 2406.11768 • Published Jun 17, 2024 • 24
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark Paper • 2410.19168 • Published Oct 24, 2024 • 24