Granite-4.1-3B-SFT-Fable5 — GGUF

Quantized GGUF builds of Granite-4.1-3B-SFT-Fable5, a ibm-granite/granite-4.1-3b model supervised-fine-tuned on the FABLE-5 Complete-2M trace corpus. These files run locally with llama.cpp, Ollama, LM Studio, and any GGUF-compatible runtime — no GPU required for the smaller quants.

Overview

Fine-tuned model Granite-4.1-3B-SFT-Fable5
Base model ibm-granite/granite-4.1-3b
Parameter class 3B
Model family dense
Training method LoRA SFT (distillation), assistant-only loss masking
Dataset FABLE-5 Complete-2M traces (private)
Format GGUF (this repo) · safetensors (merged repo)

What is FABLE-5 Complete-2M?

This model was fine-tuned on FABLE-5 Complete-2M, the full ~2M-trace FABLE-5 corpus (cleaned). Each target completion may include a <think>…</think> reasoning span followed by the response; training used assistant-only loss masking so the model learns to produce the response, not echo the prompt. The dataset is private; the fine-tuned weights are public.

Available Quantizations

File Quant Size Notes
granite-4.1-3b-sft-fable5.q4_k_m.gguf Q4_K_M ~2.1 GB Recommended — best quality/size balance
granite-4.1-3b-sft-fable5.q5_k_m.gguf Q5_K_M ~2.4 GB Higher quality
granite-4.1-3b-sft-fable5.q8_0.gguf Q8_0 ~3.6 GB Maximum quality (near-lossless)

Which to pick: Q4_K_M is the best size/quality trade-off for most users. Use Q5_K_M if you have spare RAM/VRAM and want a little more fidelity, or Q8_0 for near-lossless output when size is not a concern.

Usage

Ollama

ollama run hf.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Fable5-GGUF:Q4_K_M "Write a short story about a clockwork fox."

llama.cpp

# One-shot
llama-cli -hf ermiaazarkhalili/Granite-4.1-3B-SFT-Fable5-GGUF --jinja -p "Write a short fable about ambition." -n 512
# Interactive chat
llama-cli -hf ermiaazarkhalili/Granite-4.1-3B-SFT-Fable5-GGUF --jinja -cnv

llama-cpp-python

from llama_cpp import Llama

llm = Llama.from_pretrained(
    repo_id="ermiaazarkhalili/Granite-4.1-3B-SFT-Fable5-GGUF",
    filename="*q4_k_m.gguf",
    n_ctx=4096,
)
out = llm.create_chat_completion(
    messages=[{"role": "user", "content": "Write a short fable about ambition."}],
    max_tokens=512,
)
print(out["choices"][0]["message"]["content"])

Intended use & limitations

Research and non-commercial experimentation with FABLE-5-style creative / agentic generation. As GGUF quantizations these carry unavoidable quality loss versus the source safetensors weights — prefer Q8_0 when fidelity matters. Inherits every limitation of the base model ibm-granite/granite-4.1-3b and the source fine-tune Granite-4.1-3B-SFT-Fable5. Verify outputs before any downstream use.

Citation

@misc{granite_4_1_3b_fable5_gguf,
  author       = {Ermia Azarkhalili},
  title        = {Granite-4.1-3B-SFT-Fable5 — GGUF quantized},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Fable5-GGUF}}
}
Downloads last month
175
GGUF
Model size
3B params
Architecture
granite
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ermiaazarkhalili/Granite-4.1-3B-SFT-Fable5-GGUF

Quantized
(2)
this model