Instructions to use hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S # Run inference directly in the terminal: llama cli -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S # Run inference directly in the terminal: llama cli -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S # Run inference directly in the terminal: ./llama-cli -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
Use Docker
docker model run hf.co/hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
- LM Studio
- Jan
- vLLM
How to use hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
- Ollama
How to use hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF with Ollama:
ollama run hf.co/hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
- Unsloth Studio
How to use hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF to start chatting
- Pi
How to use hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF with Docker Model Runner:
docker model run hf.co/hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
- Lemonade
How to use hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
Run and chat with the model
lemonade run user.Kimi-K3-REAP640ja-IQ1_S-GGUF-IQ1_S
List all available models
lemonade list
- Hermes Agent
How to use hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF:IQ1_S
Run Hermes
hermes
- Atomic Chat
Kimi-K3-REAP640ja-IQ1_S-GGUF
The Japanese-calibrated sibling of Kimi-K3-REAP640-IQ1_S-GGUF: Kimi-K3 for Japanese, on a single 512 GB Mac Studio.
Same recipe as REAP640 — Unsloth's UD-IQ1_S dynamic 1-bit quant (594 GB, all 896 experts) REAP-pruned to 640 experts / 441 GB — but the keep-list comes from a Japanese + Chinese calibration instead of English + code. Same size, same expert count per layer; the two builds differ only in which 640 experts survive (~473 of 640 per layer are shared, ~167 swapped).
On ELYZA-tasks-100 (Japanese instruction benchmark, blinded LLM judge) this build scores 4.16/5 where the en+code sibling scores 1.81/5, winning the blinded pairwise comparison 83–6 with 11 ties. Japanese held-out perplexity drops from 19.46 to 4.46.
| experts | 640 of 896 per MoE layer (uniform), REAP saliency ranking |
| calibration | Japanese + Chinese subset of a tagged multi-domain corpus — 97.3% ja+zh saliency mass retained (en+code retention drops to 66.8%) |
| quantization | untouched — surviving experts are byte-identical to UD-IQ1_S (slab copy along the expert axis, no requantization) |
| router / norms | F32, inherited intact from the Unsloth quant |
| size | 441.4 GB, single file (fits 512 GB unified memory with headroom) |
| measured | ELYZA-tasks-100 generation on Mac Studio M3 Ultra 512 GB: ~3.0 tok/s effective decode |
Chinese rides along (it was 10% of the calibration and Japanese borrows heavily from Chinese experts — 42.8% top-expert overlap): zh perplexity is 4.10 vs the sibling's 7.93.
This is not a coding build. Code perplexity roughly doubles (2.00 → 3.87) and agentic use is untested here. For coding agents, use REAP640.
Full write-up (Japanese): 枝刈りKimi K3の日本語版を作って、ELYZA-tasks-100で比べた
Verified vs. not verified
Honest scorecard: exactly what has been measured, and what has not.
Verified:
| claim | evidence |
|---|---|
| Loads and serves on one 512 GB M3 Ultra, full Metal offload | 100-task ELYZA generation run end-to-end (5.8 h, ~3.0 tok/s effective) |
| Japanese generation quality vs the en+code sibling | ELYZA-tasks-100: rubric mean 4.16/5 vs 1.81/5; blinded order-randomized pairwise 83–6–11 (judge: Claude Sonnet — numbers are not comparable to GPT-4-judged leaderboards, only between these two builds) |
| Held-out perplexity (48×2048-token chunks, C4-validation-based) | ja 4.46 / zh 4.10 / en 8.50 / code 3.87 (sibling: ja 19.46 / zh 7.93 / en 7.44 / code 2.00) |
| Pruning is lossless for surviving experts | identity-prune is byte-identical (pinned by tests); router/norms stay F32 |
| Generation settings disclosed | max_tokens 4096, thinking_effort low, temp 1.0, top-p 0.95, identical for both builds; 4/100 answers hit the 4096 cap. A 16k-budget recheck rescued none of them: 2 re-ran to clean completion within the original budget (stochastic thinking runaways at temp 1.0), 1 was still empty at 16k (114k chars of thinking), 1 hit 16k again in the answer body. The cap is not the bottleneck |
Not verified:
| open question | status |
|---|---|
| Agentic use (Kimi Code CLI, tool calling) | never tested on this build — the SWE-Lancer results on the sibling's card do not transfer |
| Coding benchmarks | not measured; perplexity says expect degradation |
| Japanese factual accuracy | the judge flagged factual errors even in fluent answers; 1.6-bit experts are fluent before they are precise |
| Long-context quality | ELYZA prompts are short; 131k context is configured but unexercised here |
| Vision | mmproj not included; text tensors only |
Download & run
This repo ships one 441 GB GGUF, no shards (the Hub allows files up to 500 GB;
hf download resumes interrupted transfers):
hf download hellohazime/Kimi-K3-REAP640ja-IQ1_S-GGUF \
Kimi-K3-REAP640ja-IQ1_S.gguf --local-dir .
Kimi-K3 support is not in mainline llama.cpp yet. Build the Unsloth fork at its K3 PR (built on top of llama.cpp PR #26185):
git clone https://github.com/unslothai/llama.cpp
cd llama.cpp && git fetch origin pull/48/head:kimi-k3 && git checkout kimi-k3
cmake -B build -DGGML_METAL=ON # Apple Silicon; use -DGGML_CUDA=ON on NVIDIA
cmake --build build --config Release -j --target llama-server
./build/bin/llama-server -m Kimi-K3-REAP640ja-IQ1_S.gguf \
--port 8090 -ngl 99 -c 131072 --jinja --cache-reuse 0 \
--temp 1.0 --top-p 0.95
--cache-reuse 0is required: partial prefix-cache reuse corrupts the KDA recurrent state (known issue, see the PR discussion).- K3 is thinking-only; reasoning arrives in
reasoning_content. Control depth withchat_template_kwargs: {"thinking_effort": "low" | "high" | "max"}. - Sampling per Moonshot:
temperature 1.0, top_p 0.95.
How it was made
Identical pipeline to the sibling build, documented in
01554/kimi-k3-gguf-prune. A
tagged multi-domain calibration corpus (ja 30% / en 25% / code 25% / zh 10% /
5-language tail 10%, token shares) was streamed once through the 1.56 TB MXFP4
source with per-source saliency recording; the keep-640 plan sums only the
lang-ja + chinese labels. Saliency measurement and planning use
pipenetwork's kimi-k3-mlx
scripts (reap_calibrate.py / reap_subset.py / reap_plan.py); the GGUF
slicing (byte-slab copy along the expert axis, router rows and exp_probs_b
renumbered to keep order) is this project's only original code.
Credits: Moonshot AI (Kimi-K3), Unsloth (dynamic 1-bit quant whose protected router/norms this build inherits), Cerebras REAP (saliency criterion), kimi-k3-mlx (calibration machinery), ELYZA (ELYZA-tasks-100).
日本語の説明
Moonshot AIの2.8兆パラメータモデル Kimi-K3を、Mac Studio(512GB) 1台で動くサイズ(441GB)に枝刈りした日本語版です。
公開済みのREAP640は 英語+コードで校正したため、日本語を入れると中国語混じりの出力に崩れます。 この版は校正を日本語+中国語に変えて、同じ640個構成で作り直したものです。 サイズも構成も同じで、残っているexpertの顔ぶれだけが違います。
ELYZA-tasks-100(盲検・提示順ランダムのClaude Sonnet判定)で4.16/5。 英語+コード版は同条件で1.81/5でした。判定者がリーダーボード(GPT-4系)と 違うため、数値を外部のスコアと直接比較はできません。
注意点は2つ。コーディング能力は劣化しています(code perplexity 2.00→3.87)。 エージェント用途の検証(Kimi Code CLI、SWE-Lancer)はこの版では行っていません。 その用途にはREAP640を使ってください。
作った経緯と実測の詳細: 枝刈りKimi K3の日本語版を作って、ELYZA-tasks-100で比べた
- Downloads last month
- 73
1-bit