EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling Paper • 2610.02298 • Published 8 days ago • 57
ROWBench: Do Video Models Render What the Program Specifies? Paper • 2610.02205 • Published 8 days ago • 72
AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video Paper • 2609.14462 • Published 26 days ago • 26
Kaininja: Extending Native 3D Generators to the Part Level Paper • 2609.15659 • Published 25 days ago • 18
WorldSculpt: Generating Compositional Worlds from Grounded Videos Paper • 2609.05416 • Published Sep 4 • 26
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published Sep 3 • 50
Marionette: Predicting World States, Rendering Geometry, Painting Appearance Paper • 2608.14530 • Published Aug 14 • 35
HelloWorld: Enabling Socially Interactive Characters in Video World Models Paper • 2608.05070 • Published Aug 5 • 41
LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models Paper • 2607.08770 • Published Jul 9 • 36
BRDFusion: Physics Meets Generation for Urban Scene Inverse Rendering Paper • 2606.17049 • Published Jun 15 • 29
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models Paper • 2606.12412 • Published Jun 10 • 21