microsoft/VibeVoice-ASR-Streaming-1.5B Automatic Speech Recognition • 3B • Updated 7 days ago • 4.72k • 48
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 15 days ago • 40
Muse Glimmer Collection Muse Glimmer 30B: multimodal agentic model for local deployment. BF16 weights, GGUF k-quants, ExecuTorch builds, DFlash drafter. • 4 items • Updated Aug 10 • 107
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF Image-Text-to-Text • 27B • Updated 8 days ago • 615k • 769
Running on CPU Upgrade Featured 3.3k The Smol Training Playbook 📚 3.3k The secrets to building world-class LLMs
QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 Image-Text-to-Text • 28B • Updated 13 days ago • 30.4k • 118
Granite 4.2 Language Models Collection Efficient reasoning and thinking language models for multilingual generation, coding, and AI assistant workflows. • 3 items • Updated 16 days ago • 39
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios Paper • 2603.09983 • Published Feb 12 • 4
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution Paper • 2608.16157 • Published 25 days ago • 109