Instructions to use keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16 # Run inference directly in the terminal: llama cli -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16 # Run inference directly in the terminal: llama cli -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
Use Docker
docker model run hf.co/keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
- LM Studio
- Jan
- vLLM
How to use keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
- Ollama
How to use keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF with Ollama:
ollama run hf.co/keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
- Unsloth Studio
How to use keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF to start chatting
- Pi
How to use keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF with Docker Model Runner:
docker model run hf.co/keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
- Lemonade
How to use keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
Run and chat with the model
lemonade run user.Gemma-4-31B-it-MixQ-Q3-16G-GGUF-F16
List all available models
lemonade list
- Hermes Agent
How to use keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default keyuan01/Gemma-4-31B-it-MixQ-Q3-16G-GGUF:F16
Run Hermes
hermes
- Atomic Chat
Gemma 4 31B Mix-Quant Q3 GGUF
File
- Model:
gemma-4-31b-mixq-q3.gguf - Multimodal projector:
mmproj-gemma-4-31b-f16.gguf
What This Is
This is a conservative mixed-quant GGUF build of Gemma 4 31B for llama.cpp.
It is not a pure uniform quant. It was built with:
- importance-guided quantization using
imatrix - higher precision on more sensitive tensors
- lower precision on less sensitive tensors
Quantization Type
This release is a GGUF quantized model for llama.cpp.
Quantization family:
GGUFllama.cppimatrix-guided quantization- mixed tensor quantization (
Mix-Quant) - Q3-centered mixed recipe
This means the model is not stored with one single quant type everywhere. Instead, different tensor groups are assigned different precision levels according to sensitivity.
Importance Matrix (imatrix)
This build uses llama.cpp importance matrix calibration.
Core formula:
I_j = Σ_t x_{t,j}^2
Where:
x_{t,j}is the activation value of channeljfor token/sample steptI_jis the accumulated importance score of that channel across calibration text
Practical meaning:
- channels that activate more often and with larger magnitude get larger importance values
- more important directions are better preserved during quantization
- less important directions can be compressed more aggressively
imatrix does not use benchmark scores directly.
It estimates sensitivity from activations collected on calibration data.
Multimodal Support
Yes. Multimodal remains supported when used together with:
mmproj-gemma-4-31b-f16.gguf
Notes:
- the text model and the projector are separate files
- the text GGUF alone is not enough for vision input
- for image support, load both the main model and
mmproj
Quantization Road
The practical road was:
- Start from the original HF Gemma 4 31B model.
- Export text model to F16 GGUF.
- Build an importance matrix from calibration text.
- Use mixed quantization instead of a pure uniform Q3.
- Test with local smoke checks and benchmark samples.
Self Tests
Observed checks during the project:
- model loads successfully in
llama.cpp - dual GPU CUDA loading works
- multimodal chain remains available when
mmprojis present
Local benchmark references:
F16 MMLU-Pro 1400:0.6421428571Q3 MMLU-Pro 1400:0.6450000000Q3 HellaSwag 200:0.8650000000Q3 HellaSwag 1400:0.8771428571
Evaluation files:
evals/mmlu_pro_choose_f16_seed_20260411_n1400.jsonevals/mmlu_pro_choose_q3_seed_20260411_n1400.jsonevals/hellaswag_choose_q3_seed_20260411_n200_new.jsonevals/hellaswag_choose_q3_seed_20260411_n1400_new.json
Note:
- the strongest directly saved local F16 baseline from this project is
MMLU-Pro 1400 - HellaSwag F16 was not preserved as a finalized release reference file
Environment Build
This line was built and tested with:
- Ubuntu 20.04
- NVIDIA driver 580.95.05
- 2x RTX 5090
- CUDA 12.8 toolkit
llama.cppbuild with CUDA support
Minimal environment steps:
- Install CUDA toolkit.
- Build
llama.cppwith CUDA enabled. - Keep
mmproj-gemma-4-31b-f16.ggufnext to the main GGUF if vision is needed. - Run with
-ngl 999 -fa on.
Example text+vision server:
/home/kasm-user/src/llama.cpp-b8756/build/bin/llama-server \
-m 'gemma-4-31b-mixq-q3.gguf' \
--mmproj 'mmproj-gemma-4-31b-f16.gguf' \
-ngl 999 -fa on --ctx-size 4096 -np 1 --port 18081
Datasets Used In The Project
These datasets were used in the broader tuning and testing workflow around this model line:
TeichAI/glm-4.7-2000xFarseen0/opus-4.6-reasoning-sft-12k
Apache note:
- the project workflow treated these datasets as Apache-licensed sources
- if you redistribute publicly, verify the upstream dataset pages and their source lineage again
Practical Summary
This version is the larger and more conservative mixed-quant line.
Use this version if you want:
- stronger quality retention than the smaller 14G build
- multimodal compatibility with the same
mmproj - a safer baseline for comparison
- Downloads last month
- 105
We're not able to determine the quantization variants.