SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking Paper • 2609.13141 • Published 5 days ago • 57
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 8 days ago • 58
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 2 days ago • 263
HyQuant: Hybrid-Precision Quantization for LLM Attention Paper • 2608.27875 • Published 19 days ago • 31
Erase-then-Delta Attention: Decoupling Erase and Write Addresses in Delta-Rule Linear Attention Paper • 2606.26560 • Published Jun 25 • 3
Kalman Delta Networks: Uncertainty-aware Associative Memory Paper • 2609.07816 • Published 9 days ago • 29
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation Paper • 2609.02998 • Published 14 days ago • 20
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation Paper • 2609.05295 • Published 12 days ago • 16
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published 12 days ago • 23
MDN: Parallelizing Stepwise Momentum for Delta Linear Attention Paper • 2605.05838 • Published May 7 • 6
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 13 days ago • 184
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 13 days ago • 83
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 13 days ago • 97
Post-Training Language Models for Gold-Medal Performance in Coding Competitions Paper • 2609.02849 • Published 14 days ago • 12