DeepSeek V4 Flash 0731 — DS4 Quality128
Clean official weights, exact MXFP4 experts, maximum resident quality.
Required DS4 version: this model is not compatible with an older DS4 build. Its native MXFP4 tensors require a recent ds4 version from the main branch. That version runs both the target model and the supplied DSpark support model.
Validation status: conversion and CPU structural validation are complete. Metal generation, DSpark acceptance, throughput, peak-memory and actual one-million-token-context tests for this exact artifact are still pending.
This is a quality-first DS4 package of the official
deepseek-ai/DeepSeek-V4-Flash-0731
checkpoint. It is designed to keep the target model—and optionally its real
three-stage DSpark speculative drafter—resident on a 128 GB Apple Silicon host.
This model was quantized directly from the original official checkpoint. No behavioral weight edit was applied before quantization. All model weights were regenerated from the official local FP8 checkpoint; no tensor values were copied from another GGUF or quantized model.
Artifacts
| File | Purpose | Bytes | GiB | SHA-256 |
|---|---|---|---|---|
DeepSeek-V4-Flash-0731-DS4-Quality128.gguf |
Authoritative 43-layer target model | 102,826,238,912 |
95.7644 |
efcbf786154c2aec61511785d04a4c8d98ee42500b60ff8a810717fc0a69d1d3 |
DeepSeek-V4-Flash-0731-DS4-Quality128-DSpark-support.gguf |
Three-stage speculative drafter; not standalone | 7,297,737,120 |
6.7965 |
393b807aa0a409f1c006f0e700bbb91a80d3de28daa56c222edf944250075c66 |
| Combined GGUFs | — | 110,123,976,032 |
102.5609 |
— |
BUILD_MANIFEST.json, PROVENANCE.md and SHA256SUMS provide the full
machine-readable build record and integrity inventory.
Quantization profile
Main target model
| Tensor class | Quantization/storage | Rationale |
|---|---|---|
| Routed gate/up/down on layers 10, 14, 30, 34, 37, 38, 39, 40, 41 and 42 | Exact native MXFP4 |
Preserves source packed I8 codes and F8_E8M0 scales byte-for-byte on sensitivity-selected layers. |
| Routed gate/up on the other 33 MoE layers | IQ2_XXS, importance-matrix calibrated |
Applies the most aggressive compression to the largest tensor bank. |
| Routed down on the other 33 MoE layers | Q2_K, importance-matrix calibrated |
More conservative two-bit storage on the projection that writes expert output to the residual stream. |
| Attention projections | Q8_0 |
Protects a dense path used for every token. |
| Shared experts | Q8_0 |
Protects the expert path active for every token. |
| Vocabulary/output head | Q8_0 |
Protects final-logit fidelity. |
Indexer attn_q_b tensors on layers 2, 4, …, 42 |
forced F16 |
Preserves a small, sensitive compressed-attention component. |
| Remaining control, normalization, routing and auxiliary tensors | template-declared F16/F32/I32 |
Avoids forcing small or numerically sensitive tensors into the low-bit expert rules. |
Observed main-GGUF inventory from strict DS4 inspection:
| GGUF type | Tensors |
|---|---|
F32 |
492 |
F16 |
359 |
I32 |
3 |
Q8_0 |
345 |
IQ2_XXS |
66 |
Q2_K |
33 |
MXFP4 |
30 |
| Total | 1,328 |
The target GGUF is version 3, describes approximately 284.33 billion logical parameters and retains the checkpoint's declared 1,048,576-token training context. That declaration is not evidence that a one-million-token inference run fits or remains robust on a particular machine.
DSpark support model
This is the actual 0731 three-stage DSpark module (mtp.0–mtp.2), not the
legacy single-stage MTP attachment. The target model remains authoritative and
verifies speculative proposals.
| Parameter | Value |
|---|---|
| Stages | 3 |
| Proposal block size | 5 |
| Target layers | 40, 41, 42 |
| Markov rank | 256 |
| Noise token ID | 128799 |
| Routed gate/up | IQ2_XXS (6 tensors) |
| Routed down | exact native MXFP4 (3 tensors) |
| Dense projections | Q8_0 (31 tensors) |
| Control/auxiliary tensors | 7 F16 + 34 F32 tensors |
| Total | 81 tensors |
The support GGUF describes approximately 19.85 billion logical parameters and must be loaded alongside its matching target GGUF.
Importance calibration
The routed-expert calibration source is the 0731-native matrix from
ox-ox/DeepSeek-V4-Flash-0731-GGUF:
- Repository revision:
6d58a3a36030c3ccb969bb5759fc6ae08cd299f8 - File:
imatrix/DeepSeek-V4-Flash-0731-chat-v2-routed-moe-ds4-1p5m.dat - Size:
450,892,654bytes - SHA-256:
6fce7674df701de544e5d3351aab04e67602eddeafeb48cf70e77ebe47239eb4 - Coverage: exactly 129 routed target tensors (43 layers × gate/up/down), with exact name and vector-dimension validation
The build ran strict imatrix validation. Preserved native MXFP4 tensors do not consume imatrix data because they are copied from the official FP8 checkpoint's packed source representation rather than requantized.
The published matrix has no native DSpark entries. The support-model importer therefore made deterministic target-layer proxy aliases:
mtp.0← target layer 40mtp.1← target layer 41mtp.2← target layer 42
The extended matrix contains 138 entries and has SHA-256
689b446ed2e2657ebcb69a6516781ee0444e07641aeb6125233fb4e6fe7cbce3.
This is an explicitly recorded proxy, not a fresh activation capture from the
DSpark drafter.
Weight provenance and metadata template
The sole source of weight values was the local copy of
deepseek-ai/DeepSeek-V4-Flash-0731 at official revision
9e165c30e2704aec5d9d593cce3eebd58bbef1cb. All 48 FP8 weight shards—
166,886,535,336 bytes in total—were fully SHA-256 verified and guarded
against mutation throughout conversion.
The exact-0731 GGUF from
antirez/deepseek-v4-gguf
was used only as a bounded metadata, tokenizer, tensor-order and shape template:
- Repository revision:
1cd7b564460821938add0475a60b942c409295e0 - Template LFS SHA-256:
ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0 - Template Xet object:
7da16e1025c1b856490c29c341f4e467d15cb195389c70383dade5e6108799ac - Verified 64 MiB header SHA-256:
f0e1d5e8f3b008402aa6eb32cada3873dd926c8bf5e7d00d7788eec65f09dd6d - Template weight values used: no
Runtime requirement
Use a recent ds4 version from the main branch.
This is required because the main GGUF contains 30 native MXFP4 routed-expert
tensors. An older runtime without the MXFP4 loader and Metal kernels now in
main is not compatible, even if it can parse the GGUF header.
Install or update a main-branch checkout:
git clone https://github.com/antirez/ds4.git
cd ds4
git switch main
git pull --ff-only
make
Build note: the d516d4e quantizer PR,
tracked as antirez/ds4#642,
was needed to create the preserved-MXFP4 DSpark GGUF. Users do not need that
PR to run either supplied file.
Do not substitute antirez/ds4 main, an older DS4 binary or a generic GGUF
runtime. Container parsing alone does not demonstrate correct native MXFP4 or
mixed DSpark execution.
Running with DS4
Target-only Metal inference:
./ds4 --metal \
-m DeepSeek-V4-Flash-0731-DS4-Quality128.gguf
Greedy DSpark inference:
./ds4 --metal \
-m DeepSeek-V4-Flash-0731-DS4-Quality128.gguf \
--mtp DeepSeek-V4-Flash-0731-DS4-Quality128-DSpark-support.gguf \
--dspark --temp 0
DSpark is opt-in, accelerates generation rather than prefill, and can be neutral or slower when proposal acceptance is low. Sampled decoding does not use DSpark proposals in the pinned runtime.
Structural inspection:
./ds4 --cpu --inspect \
-m DeepSeek-V4-Flash-0731-DS4-Quality128.gguf \
--mtp DeepSeek-V4-Flash-0731-DS4-Quality128-DSpark-support.gguf \
--dspark-strict
Memory planning on a 128 GB Mac
The target machine has 128 GiB (137,438,953,472 bytes) of unified physical
memory, but the more relevant ceiling for a large Metal allocation is its
approximately 121.60 GiB recommendedMaxWorkingSetSize. The remaining
roughly 6.40 GiB is not extra model capacity; macOS, applications, drivers and
untracked transient allocations still need memory. A configuration at or below
128 GiB can therefore be unsafe when it is above—or too close to—the Metal
working-set recommendation.
Fixed weight cost
| Loaded weights | GiB | Share of 128 GiB | Share of 121.60 GiB Metal recommendation |
|---|---|---|---|
| Target model only | 95.7644 |
74.8% |
78.8% |
| Target + DSpark support | 102.5609 |
80.1% |
84.3% |
These are GGUF payload sizes, before KV cache, indexed-attention scratch,
prefill workspace, verifier state and other runtime allocations. DS4's planner
uses an approximately 97.63 GiB resident span for the main model after its
mapping/alignment accounting, rather than treating the main file size as the
entire live allocation.
Context-length scaling at prefill chunk 1,024
--ctx is the total prompt-plus-completion capacity. --prefill-chunk is the
maximum prompt microbatch DS4 processes at once; it is the relevant “batch
size” for this single-session memory calculation. A larger chunk can improve
prefill throughput, but indexed-attention scratch grows with both context length
and chunk size.
The following are conservative planning estimates. “Margin” is remaining space
under the 121.60 GiB Metal recommendation, not free system RAM.
| Context | DSpark off: total | Off: margin | DSpark on: total | On: margin |
|---|---|---|---|---|
| 4,096 | 97.77 GiB |
23.83 GiB |
104.57 GiB |
17.03 GiB |
| 32,768 | 98.05 GiB |
23.55 GiB |
104.85 GiB |
16.75 GiB |
| 131,072 | 98.99 GiB |
22.61 GiB |
105.79 GiB |
15.81 GiB |
| 262,144 | 100.24 GiB |
21.36 GiB |
107.04 GiB |
14.56 GiB |
| 524,288 | 102.75 GiB |
18.85 GiB |
109.55 GiB |
12.05 GiB |
| 1,048,576 | 107.77 GiB |
13.83 GiB |
114.56 GiB |
7.04 GiB |
At ordinary 4K–128K contexts, the model weights dominate and chunk 1,024 leaves
substantial planned margin. At 512K and especially 1M, the compressed KV and
context-by-chunk attention workspace become material. DSpark adds approximately
6.80 GiB at every context length because its support model remains resident
during prefill even though speculative decoding only accelerates generation.
Prefill batch-size effect at one-million-token context
This table holds --ctx 1048576 constant and varies --prefill-chunk. Totals
are the conservative envelope of DS4's resident-span/context estimator and the
completed release planning model. Percentages use all 128 GiB of physical RAM;
the Metal margin remains the safer operational measure.
| Prefill chunk | DSpark off total | Off: 128 GiB used | Off: Metal margin | DSpark on total | On: 128 GiB used | On: Metal margin |
|---|---|---|---|---|---|---|
| 256 | 106.20 GiB |
83.0% |
15.40 GiB |
113.00 GiB |
88.3% |
8.60 GiB |
| 512 | 106.72 GiB |
83.4% |
14.88 GiB |
113.52 GiB |
88.7% |
8.08 GiB |
| 1,024 | 107.77 GiB |
84.2% |
13.83 GiB |
114.56 GiB |
89.5% |
7.04 GiB |
| 2,048 | 110.30 GiB |
86.2% |
11.30 GiB |
117.09 GiB |
91.5% |
4.51 GiB |
| 4,096 | 116.44 GiB |
91.0% |
5.16 GiB |
123.24 GiB |
96.3% |
−1.64 GiB |
The 4,096/DSpark combination exceeds Metal's recommendation despite fitting
numerically inside 128 GiB and should not be treated as resident-safe. The
2,048/DSpark combination is also tight: its 4.51 GiB planned Metal margin can
be consumed by omitted driver and verifier peaks. Chunk 1,024 is the sensible
first DSpark experiment at 1M; target-only chunk 2,048 is the more conservative
maximum-quality starting point. Chunks 256–512 provide more margin at the cost
of more prefill iterations and likely lower prompt-processing throughput.
Target-only 1M starting point:
./ds4 --metal \
-m DeepSeek-V4-Flash-0731-DS4-Quality128.gguf \
--ctx 1048576 --prefill-chunk 2048
DSpark 1M starting point, only after unloading other large applications and measuring the target-only peak:
./ds4 --metal \
-m DeepSeek-V4-Flash-0731-DS4-Quality128.gguf \
--ctx 1048576 --prefill-chunk 1024 \
--mtp DeepSeek-V4-Flash-0731-DS4-Quality128-DSpark-support.gguf \
--dspark --temp 0
Multiple sessions and server batching
The tables describe one resident session. They do not mean that N concurrent
requests cost exactly N times the displayed total: weights and some graph
workspace are shared, while each resident session needs its own KV/context
buffers and live request state. DS4 reports the aggregate minimum context-buffer
request at server startup. For multi-user serving, start with one session,
measure the process and system peak, then raise concurrency one session at a
time; do not infer a safe concurrency count from the GGUF sizes alone.
All figures above are planning estimates, not measurements of this exact artifact. Metal load, peak-memory capture and an actual 1,048,576-token run remain required. Keep several GiB of additional operational margin, watch memory pressure rather than only Activity Monitor's process RSS, and expect other active models or large applications to invalidate the table.
Reproducibility
- Build tooling revision:
e24746463a0a3e79036dd9c0472deac6ce704f08 - Build driver SHA-256:
cb1569b909de010b4aaa2a1b1b7ed8e277bcdf41e32585f1651005b41e81567f - Build-time DS4/quantizer revision:
d516d4eeb82c454aeb2831af1b1961801d6b571b - oMLX revision:
76352ed2363e42b2146243463875756e683a640d - Profile SHA-256:
bc7f6470a0cb6208addf2ac52806d976108517119b411023da3fd91b492b218c - Build run ID:
20260801-161944-54912
The tooling worktree was intentionally dirty and is cryptographically described
in BUILD_MANIFEST.json, alongside the complete profile, source-shard hashes,
calibration/template identities and exact converter commands.
Validation status
Completed on 2026-08-01:
- focused build-driver and finalization test suites: 40/40 passed;
- every one of the 48 official source weight shards fully SHA-256 verified;
- strict routed-imatrix name and vector-dimension coverage passed;
- main GGUF: exact size, 1,328 tensors, exact name set and exact type histogram;
- DSpark GGUF: exact size, 81 tensors, three stages and target layers 40/41/42;
- strict DSpark binding: 81 tensors, 0 missing, 0 invalid and 0 metadata errors;
ds4 --cpu --inspect --dspark-strictpassed;- byte-for-byte regeneration passed for all 30 main and all three DSpark native MXFP4 tensors; and
- the original build payloads passed their recorded SHA-256 checks before
atomic publication; this README was added afterward and independently added
to
SHA256SUMS.
Still required before making runtime, quality or performance claims:
- successful Metal load and deterministic generation with DSpark disabled;
- successful Metal generation with DSpark enabled and lossless target agreement;
- proposal acceptance rate and accepted tokens per target step;
- prompt-processing and generation tokens/second;
- measured peak unified memory and practical context limits on the target host;
- an actual 1,048,576-token context run; and
- capability/perplexity comparisons against the official FP8 source.
Limitations and responsible use
Ultra-low-bit expert quantization can reduce reasoning, factuality, style fidelity and long-context robustness even when dense paths and selected experts are protected. Native MXFP4 Metal kernels and the mixed DSpark path are newer than DS4's mature Q4_K path. This package does not guarantee correctness, neutrality, safety, regulatory compliance or a particular response style. Evaluate it for the intended workload and apply appropriate access controls.
License and attribution
The upstream repository and weights are MIT licensed. This quantized derivative retains that license. Credit DeepSeek-AI for the original model, ox-ox for the 0731 routed-expert importance matrix, antirez for DS4 and the exact-0731 metadata recipe, and apetersson for the DS4 fork, conversion profile and release tooling.
- Downloads last month
- 935
We're not able to determine the quantization variants.
Model tree for apetersson/DeepSeek-V4-Flash-0731-DS4-Quality128
Base model
deepseek-ai/DeepSeek-V4-Flash-0731