Qwen3.6-35B-A3B-Fast-NVFP4-GGUF

GGUF NVFP4 quantization of unsloth/Qwen3.6-35B-A3B-NVFP4-Fast, a 35B parameter MoE model with 3B active parameters.

What is the "Fast" Variant?

Unsloth's NVFP4 Fast variant is a speed-optimized NVFP4 quantization that delivers 1.79x faster throughput than other NVFP4 quants while maintaining competitive accuracy:

Variant MMLU-Pro GPQA AIME 2025
Unsloth NVFP4 Fast 85.58 87.75 91.67
Unsloth NVFP4 85.85 86.74 92.29
NVIDIA NVFP4 85.60 87.12 91.88
BF16 85.75 86.36 92.50

The Fast variant is calibrated on a mixture of Unsloth's dataset + UltraChat dataset, optimized for throughput on vLLM's native NVFP4 backend (cute-DSL/CUTLASS/flashinfer_trtllm).

About the Model

Qwen3.6-35B-A3B is a multimodal MoE model from Alibaba's Qwen team:

  • 35B total parameters, 3B active per token (256 experts, 8 active)
  • 40-layer decoder with Gated DeltaNet + full attention hybrid
  • 27-layer vision encoder (SigLIP-based) for image/video understanding
  • 262K native context (extensible to 1M+ via YaRN)
  • Multi-Token Prediction (MTP) for faster speculative decoding
  • Agentic coding with SWE-bench Verified 73.4, tool calling support

Files

File Size Description
qwen36-35b-a3b-fast-nvfp4.gguf ~18.8 GB NVFP4 quantized text model
mmproj-qwen36-35b-a3b-f16.gguf ~0.84 GB Vision encoder (F16)

Usage

llama.cpp

llama-server \
  -m qwen36-35b-a3b-fast-nvfp4.gguf \
  --mmproj mmproj-qwen36-35b-a3b-f16.gguf \
  -ngl 99 \
  --host 0.0.0.0 \
  --port 8080

vLLM (Recommended for Max Performance)

pip install vllm flashinfer-python nvidia-cutlass-dsl
vllm serve FreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF --dtype auto --max-model-len 4096

Hardware Requirements

  • Minimum: 24 GB VRAM for partial offload
  • Recommended: 32+ GB VRAM for full GPU offload
  • Optimal: Blackwell B200 for max NVFP4 throughput

Quantization

Quantized from Qwen/Qwen3.6-35B-A3B BF16 weights using llama.cpp (llama-quantize.exe NVFP4).

License

Apache 2.0 - same as the base model.

Credits

Downloads last month
5,091
GGUF
Model size
36B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF

Quantized
(3)
this model