kshitijthakkar/Kirigami-Qwen3.6-20B-A3B-NVFP4
Text Generation • 14B • Updated • 70 • 1
Qwen3.6-35B-A3B carved by expert importance to fit one consumer GPU. No training, no calibration. 800 tok/s on a laptop 5090.
How we carved a 35B MoE to fit a 24GB GPU — zero training