MiniMax-H3 · SVDQuant NVFP4 (rank 32, RTN) — reference format

The RTX 50-series (sm_120) sibling of the int4 release: e2m1 4-bit residual + bf16 rank-32 low-rank branch, quantized fresh from BF16 (int4 and fp4 grids do not nest, so transcoding is never used).

Status: reference (unpacked) format. Runs through svdquant's reference backend today; the packed fp4 export for the Blackwell tensor-core kernels is in progress. RTN rounding (fp4-GPTQ pending). Same license inheritance and credits as the int4 release.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ModelsLab/MiniMax-H3-svdquant-nvfp4_r32

Finetuned
(75)
this model