DeepSeek-V4-Flash-0731 EXL3 3.0 bpw

This repository contains a 3.0 bpw EXL3 quantization of deepseek-ai/DeepSeek-V4-Flash-0731.

Status

Gate Status
Quantization and artifact assembly Complete
Structural verification and per-file SHA-256 manifest Passed
H200 model load and CUDA kernel compilation Reached
End-to-end generation on H200 Not yet passed

The weights are complete and structurally verified, but this release is not being represented as runtime-validated. The pinned experimental H200/vLLM integration loaded the model and compiled the native SM90 CUDA path, then failed during TileLang host-DSO export with Target triple should not be empty. No successful generation result is claimed.

Artifact

  • Target routed-weight rate: 3.0 bpw
  • Achieved routed-weight rate: 3.0 bpw
  • Quantization: EXL3, mcg codebook
  • Source representation: packed E2M1 FP4 with UE8M0 scales
  • Layout: rank-sliced TP4 expert weights
  • Weight files: 177
  • Weight bytes: 124,867,114,600 (about 116.29 GiB)
  • Tensor records: 534,653

EXL3_MANIFEST.json records the complete artifact inventory and per-file SHA-256 checksums. CALIBRATION_COVERAGE.json records the calibration composition and coverage evidence.

Calibration

The quantization used 1,146,665 calibration tokens spanning four axes:

  • general language
  • legal text
  • code and agentic tasks
  • reasoning and termination behavior

Routing was observed naturally; experts were not artificially forced active. The preserved capture contained 688 files totaling 104,303,431,496 bytes, with digest e31f665d1ffc24b9f80e90ff31e804cf4336efa25f30a033c065afbcf3df5968.

Compatibility

These are experimental rank-sliced DeepSeek V4 EXL3 weights. A compatible ExLlamaV3/EXL3 implementation that understands this layout is required. Mainstream drop-in compatibility is not claimed, and users should validate generation and structured output behavior in their own target runtime before production use.

Provenance and acknowledgements

  • Base model: DeepSeek-V4-Flash-0731 by DeepSeek
  • Quantization format and runtime work: ExLlamaV3 by TurboDerp and contributors
  • H200 compute for this campaign: RunPod
  • Thanks to Inkling and JarvisLabs for support and prior infrastructure used across the broader quantization campaign

The upstream model's license, acceptable-use terms, and limitations continue to apply. This repository changes the weight representation; it does not change the underlying model license.

Downloads last month
38
Safetensors
Model size
70B params
Tensor type
I64
·
F32
·
BF16
·
F8_E4M3
·
F16
·
I16
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 0xSero/DeepSeek-V4-Flash-0731-EXL3-3.0bpw

Quantized
(70)
this model