Nemotron-Labs-3-Puzzle-75B-A9B — MLX 6-bit

MLX 6-bit affine quantization of nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16, converted for Apple Silicon.

Runtime setup

Stock mlx-lm (0.31.x) can't load this model yet — it crashes with uniform(): incompatible function arguments because NVIDIA's Puzzle architecture uses different MoE dims per layer, and mlx-lm assumes they're all the same.

About the model

Puzzle-75B-A9B is NVIDIA's deployment-optimized compression of Nemotron-3-Super-120B-A12B. It's a hybrid Mamba-2 / Attention / LatentMoE architecture (88 backbone layers) with heterogeneous per-layer expert configs - MoE intermediate sizes vary 1280–2688 and active experts per token vary 4–22 across layers. 75.3B total / 9.3B active parameters.

This conversion

  • Source: nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16
  • Format: MLX affine 6-bit, group size 64 (~6.5 bpw)
  • Size on disk: ~57 GB
  • Converted with: mlx-lm 0.31.2 + mlx 0.31.1 (CUDA backend on a Blackwell), plus local patches to nemotron_h.py to support Puzzle's heterogeneous per-layer MoE dims, LatentMoE fc1_latent_proj / fc2_latent_proj, and the model.backbone. prefix in NVIDIA's checkpoints.
  • MTP weights: mtp.safetensors (5.9 GB) from the source is included in this repo but is not currently used at inference time — mlx-lm has no Nemotron-H MTP path yet (tracking mlx-lm#1161). The tensors are preserved here so they'll be available whenever native speculative-decoding support lands.

License

Governed by the NVIDIA Open Model License. Derivative of NVIDIA's Nemotron-3 family. "Nemotron" is a trademark of NVIDIA Corporation. Not affiliated with or endorsed by NVIDIA.

Downloads last month
979
Safetensors
Model size
75B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for georgeis55/Nemotron-Labs-3-Puzzle-75B-A9B-MLX-6bit

Collection including georgeis55/Nemotron-Labs-3-Puzzle-75B-A9B-MLX-6bit