How to use from
OpenClaw
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "mlx-community/NVIDIA-Nemotron-3-Super-120B-A12B-OptiQ-2bit"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "mlx-community/NVIDIA-Nemotron-3-Super-120B-A12B-OptiQ-2bit" \
  --custom-provider-id mlx-lm \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

mlx-community/NVIDIA-Nemotron-3-Super-120B-A12B-OptiQ-2bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. Read the write-up · All OptiQ quants · Docs

A 120-billion-parameter model that runs on a 36 GB Mac. This is a 2-bit mixed-precision MLX quant of NVIDIA's Nemotron-3-Super-120B-A12B (247 GB at bf16), produced by mlx-optiq. It is 47.5 GB on disk. While it generates, only ~14 GB sits in RAM: the Mamba blocks, attention, router and shared experts stay resident, and the 34 GB of routed mixture-of-experts weights stream off the SSD as the router selects them, through optiq serve --stream-experts.

Nemotron-3-Super is a hybrid: Mamba2 state-space blocks interleaved with attention and a 512-expert sparse MoE (22 active per token). Asked to write Flappy Bird in a single HTML file, the 2-bit model produced a complete, working game. Here it is playing it:

Nemotron-3-Super-120B-A12B 2-bit playing the Flappy Bird it wrote, on a 36 GB Mac

What it is

Property Value
Base NVIDIA-Nemotron-3-Super-120B-A12B (hybrid Mamba2 + attention + 512-expert MoE, 22 active)
Method OptiQ static — structural per-layer bit allocation, no calibration
Bit-widths 4-bit on Mamba / attention / router / shared experts / edges, 2-bit on the routed experts
Achieved bits-per-weight 2.52
On disk 47.5 GB
Resident while running ~14 GB (routed experts streamed)
Decode speed ~3 tok/s on an M3 Max (36 GB)

For a model this large, exact calibration-driven sensitivity is impractical (it would run for days and needs the full model resident as a reference), so OptiQ's static method assigns bits from architecture alone. See the methods comparison.

Run it

This is a Nemotron hybrid (model_type: nemotron_h), so it needs mlx-lm from main and import optiq (install from git, not a version pin):

pip install -U mlx-optiq "mlx-lm @ git+https://github.com/ml-explore/mlx-lm.git"

Serve it with SSD expert streaming (auto-enabled for a MoE too big to fit resident; --stream-experts forces it):

optiq serve --model mlx-community/NVIDIA-Nemotron-3-Super-120B-A12B-OptiQ-2bit --stream-experts

Then open the Lab, ask for a game, and watch it render in the Canvas pane. Only the routed experts stream per token; the Mamba state, attention and shared experts stay resident, so the footprint stays ~14 GB no matter how large the model on disk is.

Notes

This is an extreme quant. 2-bit on the routed experts is lossy, and the point of this artifact is that a 120 B hybrid MoE runs at all on consumer Apple Silicon, with coherent output. For reference quality on this base, use the bf16 weights or a higher-bit quant. The full story is in the blog post.

Downloads last month
1,232
Safetensors
Model size
121B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/NVIDIA-Nemotron-3-Super-120B-A12B-OptiQ-2bit

Quantized
(56)
this model