Qwen3-4B-Instruct-2507 GPTQ INT4/G256 Safetensors

Pre-LLiMa Hugging Face checkpoint for Sima.ai compilation. Source revision was not captured locally; pin an immutable revision before publication.

Quantization

  • All decoder Linear layers and lm_head: GPTQ, symmetric INT4, group size 256, static actorder.
  • Calibration: HuggingFaceH4/ultrachat_200k, 512 samples, 1024 tokens, batch size 1.
  • No intentional mixed-precision language layers.

Full WikiText evaluation

EleutherAI/wikitext_document_level, wikitext-2-raw-v1, full run (2026-07-16).

Checkpoint Word perplexity
Qwen/Qwen3-4B-Instruct-2507 14.271144
This GPTQ checkpoint 15.294486

Reproduce

python quantize.py --model-path /project/mlasw/share/huggingface/models--Qwen--Qwen3-4B-Instruct-2507 --output-dir /path/to/output

Included: quantized safetensors, tokenizer/configuration, quantize.py, recipe.yaml, and versions.txt.

Validation and limitations

Saved scales were checked for finite values. Transformers load/generation smoke test is pending before release. Quantization quality can vary by language, prompt format, domain, and context length; validate the compiled deployment on target hardware.

Downloads last month
51
Safetensors
Model size
4B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for simaai/Qwen3-4B-Instruct-2507-GPTQ-Safetensors

Quantized
(276)
this model

Collection including simaai/Qwen3-4B-Instruct-2507-GPTQ-Safetensors