Gemma-4-12B-it-Uncensored-Heretic-NVFP4-GGUF

NVFP4 GGUF quantization of llmfan46/gemma-4-12B-it-uncensored-heretic - an uncensored/heretic (abliterated) finetune of Google's Gemma 4 12B with vision support.

About NVFP4

NVFP4 is NVIDIA's native 4-bit floating point format (E4M3) designed for Blackwell architecture GPUs (RTX 50-series, B100/B200). It provides:

  • Native tensor core acceleration on Blackwell GPUs
  • Better dynamic range than INT4 formats due to floating point representation
  • No dequantization overhead - processed directly in FP4

When to use NVFP4 vs other formats:

  • NVFP4 - Best for Blackwell GPUs (RTX 5060 Ti, 5070, 5080, 5090, B100, B200)
  • Q4_K_M - Best for pre-Blackwell GPUs and CPU inference
  • MXFP4 - Open standard, works on any GPU with MX support

Files

File Type Size Description
gemma4-12b-heretic-nvfp4.gguf NVFP4 ~6.5 GB Text model (4.68 BPW)
mmproj-gemma-4-12b-heretic-f16.gguf F16 ~116 MB Vision encoder (mmproj)

Quantization Details

Property Value
Format NVFP4 (E4M3)
Bits Per Weight 4.68 BPW
Source Model llmfan46/gemma-4-12B-it-uncensored-heretic
Architecture Gemma4UnifiedForConditionalGeneration
Layers 48
Hidden Size 3840
Context Length 262144
Vision Yes (Gemma4V projector)
Thinking Enabled by default (opt-out via enable_thinking=false)

Model Description

This is an abliterated (uncensored/heretic) finetune of Google's Gemma 4 12B, a multimodal model with both text and vision capabilities. The original model was finetuned to remove safety alignment restrictions while maintaining the model's core capabilities.

Gemma 4 features a hybrid attention architecture with alternating sliding window and full attention layers, native vision encoding, and tool calling support.

Usage

llama.cpp CLI

# Text only
./llama-cli -m gemma4-12b-heretic-nvfp4.gguf -p "Hello" -n 100

# With vision (requires mmproj)
./llama-server -m gemma4-12b-heretic-nvfp4.gguf \
  --mmproj mmproj-gemma-4-12b-heretic-f16.gguf \
  --host 0.0.0.0 --port 8080 -ngl 99

LM Studio

  1. Download both files
  2. Load gemma4-12b-heretic-nvfp4.gguf as the model
  3. Load mmproj-gemma-4-12b-heretic-f16.gguf as the mmproj
  4. The model supports image inputs via the vision encoder

huggingface-hub

from huggingface_hub import hf_hub_download

model_path = hf_hub_download(
    repo_id="FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-NVFP4-GGUF",
    filename="gemma4-12b-heretic-nvfp4.gguf"
)
mmproj_path = hf_hub_download(
    repo_id="FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-NVFP4-GGUF",
    filename="mmproj-gemma-4-12b-heretic-f16.gguf"
)

Quantization Pipeline

  1. Download source: llmfan46/gemma-4-12B-it-uncensored-heretic
  2. Convert to F16 GGUF: convert_hf_to_gguf.py --outtype f16
  3. Extract mmproj: convert_hf_to_gguf.py --mmproj --outtype f16
  4. Quantize text: llama-quantize input-f16.gguf output-nvfp4.gguf NVFP4

Hardware Requirements

Component Requirement
GPU NVIDIA Blackwell (RTX 50-series) for full acceleration
VRAM ~7 GB minimum
RAM ~16 GB recommended
Storage ~7 GB

License

Apache 2.0 (inherited from base model)

Downloads last month
2,064
GGUF
Model size
12B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-NVFP4-GGUF

Quantized
(9)
this model