Qwen3-VL-8B ChartQA — GGUF Q4_K_M

Persistent CPU-demo artifact converted from the pinned fine-tuned merged checkpoint steven0226/qwen3vl-8b-chartqa-merged-16bit@519060ef43df3261e0512e5ae4c82a4d4e675f32.

Files

  • Qwen3VL-8B-ChartQA-Q4_K_M.gguf — fine-tuned language model, Q4_K_M
  • mmproj-Qwen3VL-8B-ChartQA-Q8_0.gguf — fine-tuned vision encoder/projector, Q8_0

Conversion used llama.cpp commit 79bba02a6741de194912d370015866414faa83ad. conversion_metadata.json records byte sizes, SHA-256 digests and the raw ChartQA smoke verification.

Verification scope

Raw ChartQA test row 41 (without the presentation overlay used in portfolio case images) returned the expected answer 96. This is a smoke verification, not a complete GGUF quality evaluation.

The formal 2,500-question quality result belongs to the separately evaluated AWQ/vLLM artifact: 85.52%, -0.72 pp versus merged 16-bit, passing the predefined 2 pp gate. Do not treat that number as a GGUF score.

llama.cpp

Use a recent llama.cpp build with Qwen3-VL support:

llama-server \
  --model Qwen3VL-8B-ChartQA-Q4_K_M.gguf \
  --mmproj mmproj-Qwen3VL-8B-ChartQA-Q8_0.gguf \
  --ctx-size 4096 --jinja

中文摘要

這是 ChartQA fine-tuned Qwen3-VL-8B 的持久展示用 GGUF。語言模型採 Q4_K_M,vision encoder/projector 採 Q8_0。已用不含答案標註的 raw ChartQA 圖表完成單題 smoke verification;正式品質數字仍以 AWQ/vLLM 的完整 2,500 題評估為準,不跨 inference stack 混用。

Downloads last month
283
GGUF
Model size
8B params
Architecture
qwen3vl
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for steven0226/qwen3vl-8b-chartqa-gguf

Dataset used to train steven0226/qwen3vl-8b-chartqa-gguf

Space using steven0226/qwen3vl-8b-chartqa-gguf 1