Leading the System One Mosaic Benchmark: What Darwin-27B-ZTC-v2's #1 Means FINAL-Bench • about 7 hours ago • 6
Testing selected-token J-lens directions and a fixed-symbol autoencoder on Qwen3-8B kotlarmilos • about 8 hours ago
Introducing Superfluid: the local-AI native LLM server for multi-agent inference basecompute • about 20 hours ago • 3
JOSIE-2 Technical Report: Curated Post-Training for Reasoning and Behavior in Small Language Models Goekdeniz-Guelmez • about 22 hours ago • 1
LightOnOCR-3: High-Performance OCR and Layout Extraction in One Model lightonai • about 24 hours ago • 24
GLiNER Can Now Detect Names and PII in Indian Languages, in Their Own Scripts nishikantmandal007 • 1 day ago • 3
RLCD: Are We Moving from AI Models That Generate Answers to Models That Make Calibrated Decisions? javadtaghia • 1 day ago
Claude Opus 5.5 is the new #1 on Multilingual Terminal-Bench, the coding benchmark on AURORA, our multilingual AI leaderboard. Lilt-org • 1 day ago