WebWorld: The Browser as a World Model for Self-Improving Web Code Paper • 2608.30530 • Published 10 days ago • 10
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 10 days ago • 39
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 9 days ago • 263
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published 15 days ago • 179
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published 17 days ago • 34
ClawBench & Agentic Web Benchmarks Collection ClawBench and related papers and open datasets for real-world web and computer-use agent evaluation. • 359 items • Updated 10 days ago • 2
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published Jul 14 • 108
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Paper • 2607.05382 • Published Jul 9 • 88
Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion Paper • 2606.14885 • Published Jun 12 • 13
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence Paper • 2605.26340 • Published May 25 • 38
Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback Paper • 2606.06113 • Published Jun 4 • 16
MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection Paper • 2605.30288 • Published May 29 • 23
ClawBench — Browser Agent Benchmark Suite Collection Benchmark dataset (V1+V2), live leaderboard Space, and full V1 execution traces — everything you need to run, regrade, or compare on ClawBench. • 5 items • Updated May 12 • 1
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports Paper • 2510.02190 • Published Jan 29 • 20
Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction Paper • 2605.05242 • Published May 3 • 127