Abiray commited on
Commit
8a22290
·
verified ·
1 Parent(s): 6883e73

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +68 -0
README.md ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: gguf
4
+ pipeline_tag: image-text-to-text
5
+ base_model: ATH-MaaS/OvisOCR2
6
+ tags:
7
+ - ocr
8
+ - document-parsing
9
+ - multimodal
10
+ - markdown
11
+ - tables
12
+ - formulas
13
+ - gguf
14
+ - llama-cpp
15
+ ---
16
+
17
+ # OvisOCR2 - GGUF Quantizations
18
+
19
+ <p align="center">
20
+ <img src="https://cdn-uploads.huggingface.co/production/uploads/658a8a837959448ef5500ce5/vRCIu5QD8VuIJolkC_ZHQ.png" alt="Ovis" width="30%" />
21
+ </p>
22
+
23
+ This repository contains GGUF format quantizations of **OvisOCR2**, a compact 0.8B end-to-end model for page-level document parsing. The original model was developed by **ATH-MaaS** by post-training `Qwen3.5-0.8B` to parse full document pages directly into clean Markdown (including LaTeX formulas, HTML tables, and layout components).
24
+
25
+ OvisOCR2 establishes a new state-of-the-art for compact document understanding, scoring **96.58** on OmniDocBench v1.6 and outperforming traditional, multi-stage layout analysis pipelines.
26
+
27
+ ---
28
+
29
+ ## Available Files
30
+
31
+ ### Main Text Models
32
+
33
+ | File Name | Precision / Quantization | File Size | Description |
34
+ | :--- | :--- | :--- | :--- |
35
+ | `OvisOCR2-F16.gguf` | 16-bit Float | 1.52 GB | Baseline unquantized model |
36
+ | `OvisOCR2-BF16.gguf` | 16-bit Brain Float | 1.52 GB | Native weight precision |
37
+ | `OvisOCR2-Q8_0.gguf` | 8-bit | 812 MB | Near-identical precision to F16 |
38
+ | `OvisOCR2-Q6_K.gguf` | 6-bit | 630 MB | Excellent balance of size and accuracy |
39
+ | `OvisOCR2-Q5_K_M.gguf` | 5-bit (Medium) | 578 MB | Recommended for low-resource deployment |
40
+ | `OvisOCR2-Q5_K_S.gguf` | 5-bit (Small) | 564 MB | Highly optimized 5-bit layout |
41
+ | `OvisOCR2-Q4_K_M.gguf` | 4-bit (Medium) | 529 MB | Standard 4-bit quantization |
42
+ | `OvisOCR2-Q4_K_S.gguf` | 4-bit (Small) | 505 MB | Lightweight 4-bit footprint |
43
+ | `OvisOCR2-Q3_K_M.gguf` | 3-bit (Medium) | 466 MB | Maximum compression ratio |
44
+
45
+ ### Multimodal Projectors (`mmproj`)
46
+ *Note: Because OvisOCR2 is a vision-language model, you **must** download one of these image processing units alongside your choice of the text models listed above.*
47
+
48
+ * `mmproj-F32.gguf` (402 MB) - Unquantized full precision projector.
49
+ * `mmproj-F16.gguf` (205 MB) - Recommended standard performance/size option.
50
+ * `mmproj-BF16.gguf` (207 MB) - Target alternative precision layout.
51
+
52
+ ---
53
+
54
+ ## Inference Guide (`llama.cpp`)
55
+
56
+ To run multimodal OCR tasks using these GGUF files, you need to use the `llama-minicpmv-cli` or `llama-llava-cli` tool (depending on your build version of `llama.cpp`) to handle simultaneous image and text tokens.
57
+
58
+ ### Basic Command Line Example
59
+
60
+ ```bash
61
+ # Run parsing via llama.cpp cli tools
62
+ ./llama-minicpmv-cli \
63
+ -m OvisOCR2-Q5_K_M.gguf \
64
+ --mmproj mmproj-F16.gguf \
65
+ --image /path/to/your/document_page.jpg \
66
+ -p "<|im_start|>user\nExtract all readable content from the image in natural human reading order and output the result as a single Markdown document. Format formulas as LaTeX. Format tables as HTML: <table>...</table>. Preserve the original text without translation.<|im_end|>\n<|im_start|>assistant\n" \
67
+ -n 4096 \
68
+ --temp 0.0