LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents Paper • 2609.13287 • Published 9 days ago • 15
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 10 days ago • 218
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published 15 days ago • 236
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 18 days ago • 157
CogEvol: Towards Efficient and Reliable Learning Environment Generation Paper • 2608.30968 • Published 18 days ago • 32
Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision Paper • 2608.16812 • Published Aug 17 • 50
Progressive Agent Skill Generation via Reinforcement Learning Paper • 2608.01678 • Published Aug 3 • 62
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model Paper • 2604.20796 • Published Apr 22 • 244
Innovator-VL: A Multimodal Large Language Model for Scientific Discovery Paper • 2601.19325 • Published Jan 27 • 85
AI for Service: Proactive Assistance with AI Glasses Paper • 2510.14359 • Published Oct 16, 2025 • 77
ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution Paper • 2510.12793 • Published Oct 14, 2025 • 5
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Paper • 2508.18265 • Published Aug 25, 2025 • 226