SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? Paper • 2609.09113 • Published 10 days ago • 20
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs Paper • 2603.05890 • Published Mar 6 • 93
LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth Paper • 2602.07962 • Published Feb 8 • 26