Qwen3.6-35B-A3B-Fast-MXFP4-MOE-GGUF

GGUF MXFP4 MoE quantization of unsloth/Qwen3.6-35B-A3B-NVFP4-Fast, a 35B parameter MoE model with 3B active parameters.

What is the "Fast" Variant?

Unsloth's NVFP4 Fast variant is a speed-optimized quantization that delivers 1.79x faster throughput than other NVFP4 quants. This GGUF extends that optimization to MXFP4 MoE format:

Variant MMLU-Pro GPQA AIME 2025
Unsloth NVFP4 Fast 85.58 87.75 91.67
Unsloth NVFP4 85.85 86.74 92.29
NVIDIA NVFP4 85.60 87.12 91.88
BF16 85.75 86.36 92.50

MXFP4 MoE Format

This quantization uses a hybrid approach for optimal quality:

  • Expert weights: MXFP4 (E2M1 microscaling, 4-bit)
  • Non-expert weights (attention, embeddings, norms): Q8_0 (8-bit)

MXFP4 is an open standard supported by AMD, NVIDIA, and Microsoft, making it compatible with a wider range of hardware.

About the Model

Qwen3.6-35B-A3B is a multimodal MoE model from Alibaba's Qwen team:

  • 35B total parameters, 3B active per token (256 experts, 8 active)
  • 40-layer decoder with Gated DeltaNet + full attention hybrid
  • 27-layer vision encoder (SigLIP-based) for image/video understanding
  • 262K native context (extensible to 1M+ via YaRN)
  • Multi-Token Prediction (MTP) for faster speculative decoding
  • Agentic coding with SWE-bench Verified 73.4, tool calling support

Files

File Size Description
qwen36-35b-a3b-fast-mxfp4_moe.gguf ~18.9 GB MXFP4 MoE quantized text model
mmproj-qwen36-35b-a3b-f16.gguf ~0.84 GB Vision encoder (F16)

Usage

llama.cpp

llama-server \
  -m qwen36-35b-a3b-fast-mxfp4_moe.gguf \
  --mmproj mmproj-qwen36-35b-a3b-f16.gguf \
  -ngl 99 \
  --host 0.0.0.0 \
  --port 8080

Hardware Requirements

  • Minimum: 24 GB VRAM for partial offload
  • Recommended: 32+ GB VRAM for full GPU offload

Quantization

Quantized from Qwen/Qwen3.6-35B-A3B BF16 weights using llama.cpp (llama-quantize.exe --allow-requantize MXFP4_MOE).

License

Apache 2.0 - same as the base model.

Credits

Downloads last month
2,047
GGUF
Model size
36B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for FreedomAISVR/Qwen3.6-35B-A3B-MXFP4-MOE-Fast-GGUF

Quantized
(3)
this model