chanasia commited on
Commit
38f1f1a
·
verified ·
1 Parent(s): 6f8fc53

Add model card

Browse files
Files changed (1) hide show
  1. README.md +82 -0
README.md ADDED
@@ -0,0 +1,82 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: typhoon-ai/typhoon-ocr1.5-2b
3
+ base_model_relation: quantized
4
+ license: apache-2.0
5
+ language:
6
+ - en
7
+ - th
8
+ pipeline_tag: image-text-to-text
9
+ library_name: llama.cpp
10
+ tags:
11
+ - gguf
12
+ - imatrix
13
+ - OCR
14
+ - vision-language
15
+ - document-understanding
16
+ - multimodal
17
+ - llama.cpp
18
+ ---
19
+
20
+ # typhoon-ocr1.5-2b GGUF
21
+
22
+ GGUF quantization of [typhoon-ai/typhoon-ocr1.5-2b](https://huggingface.co/typhoon-ai/typhoon-ocr1.5-2b),
23
+ a Thai/English document-OCR vision-language model built on Qwen3-VL-2B-Instruct.
24
+
25
+ This is a **vision-language model**: you need both the model file and the `mmproj`
26
+ (multimodal projector) file. The model file alone will load, but it will not see images.
27
+
28
+ ## Files
29
+
30
+ | File | Size | Notes |
31
+ | --- | --- | --- |
32
+ | `typhoon-ocr1.5-2b-Q4_K_M-imat.gguf` | 1.03 GiB | Language model, Q4_K_M with importance matrix |
33
+ | `typhoon-ocr1.5-2b-mmproj-Q8_0.gguf` | 424 MiB | Vision encoder / projector — **required** |
34
+
35
+ The Q4_K_M weights were quantized with an importance matrix (imatrix) computed from a
36
+ calibration set, which recovers some of the quality lost at 4 bits compared to a plain
37
+ Q4_K_M of the same size.
38
+
39
+ ## Usage
40
+
41
+ ### llama-server (OpenAI-compatible API)
42
+
43
+ ```bash
44
+ llama-server \
45
+ -m typhoon-ocr1.5-2b-Q4_K_M-imat.gguf \
46
+ --mmproj typhoon-ocr1.5-2b-mmproj-Q8_0.gguf \
47
+ -c 8192 --host 0.0.0.0 --port 8080
48
+ ```
49
+
50
+ Then post an image to `/v1/chat/completions` the usual way:
51
+
52
+ ```bash
53
+ curl http://localhost:8080/v1/chat/completions \
54
+ -H 'Content-Type: application/json' \
55
+ -d '{
56
+ "messages": [{
57
+ "role": "user",
58
+ "content": [
59
+ {"type": "image_url", "image_url": {"url": "data:image/png;base64,<BASE64>"}},
60
+ {"type": "text", "text": "Extract all text from this document as Markdown."}
61
+ ]
62
+ }]
63
+ }'
64
+ ```
65
+
66
+ ### llama-mtmd-cli (one-shot)
67
+
68
+ ```bash
69
+ llama-mtmd-cli \
70
+ -m typhoon-ocr1.5-2b-Q4_K_M-imat.gguf \
71
+ --mmproj typhoon-ocr1.5-2b-mmproj-Q8_0.gguf \
72
+ --image page.png \
73
+ -p "Extract all text from this document as Markdown."
74
+ ```
75
+
76
+ Use a recent llama.cpp build — Qwen3-VL support landed relatively late.
77
+
78
+ ## License
79
+
80
+ Apache 2.0, inherited from the base model. See the
81
+ [base model card](https://huggingface.co/typhoon-ai/typhoon-ocr1.5-2b) for the model's
82
+ intended use and limitations.