Typhoon-OCR 1.5 2B โ€” ROCmFP4 GGUF (Thai iMatrix Calibrated)

This repository provides ROCmFP4 quantized GGUF weights for typhoon-ai/typhoon-ocr1.5-2b, calibrated specifically for high-fidelity Thai document OCR using Importance Matrix (iMatrix).

The text backbone is quantized to Q4_0_ROCMFP4 (UE4M3-scale experimental with Q6_K token embeddings), while the Vision Projector (mmproj-f16.gguf) remains unquantized in full FP16 to ensure zero loss in visual document resolution.


๐Ÿš€ UPDATE (Sep 2026): 100% Native ROCm / HIP GPU Acceleration is LIVE!

Full GPU acceleration (Vision Transformer + LLM Backbone) is now supported natively via PR #112 (charlie12345/ROCmFPX#112)!

  • No More CPU Fallback: You no longer need --no-mmproj-offload! Both the vision encoder and LLM run 100% on the AMD GPU.
  • Blazing Speeds on AMD Radeon 890M (Strix Point):
    • โšก Vision Prompt Processing: 327.8 tok/s (entire high-res document ingested in ~6โ€“7 seconds)
    • ๐Ÿš€ Text Generation (OCR Decode): 46.7 tok/s
    • ๐Ÿ’พ Ultra-Low VRAM: Total memory footprint is only ~2.7 GB (allowing large batches or multi-turn OCR on 16GBโ€“32GB unified memory laptops).
  • Build Requirement: Until PR #112 is merged into upstream master releases, you must build the engine from source using PR #112 or the branch nanashi66:fix-cublas-f16-pointer-mismatch. See the Build Instructions below.
  • Alternative (Plug-and-Play): If you prefer not to build from source, you can still use the pre-built Vulkan backend of ROCmFPX which runs out of the box.

๐ŸŒŸ Key Highlights

  • 41% Smaller than Q8_0 (67% Smaller than FP16): Text backbone compressed down to 1.08 GB (from 3.28 GB FP16).
  • High-Precision Thai Preservation: Calibrated on a custom Thai Calibration Mix (70% general Thai corpus + 30% complex Thai official & legal documents), preventing degradation of Thai vowels, tone marks (เธงเธฃเธฃเธ“เธขเธธเธเธ•เนŒ), and specialized vocabulary.
  • 100% Empirical Document Fidelity: Verified on complex Thai government and legal documents with exact table preservation (<table>...</table>), zero character hallucinations, and 100% identical Thai numerals.
  • Tuned for AMD RDNA 3 / 3.5 (gfx1150 / gfx1100 / gfx1103): High inference speed (46.7 tok/s) and ultra-low VRAM footprint (~2.7 GB) on AMD Ryzen AI 9 HX 370 / Radeon 890M / 880M / 780M iGPUs.

๐Ÿ“Š Benchmark & Quality Verification (Measured on AMD Radeon 890M / gfx1150)

Hardware: AMD Ryzen AI 9 HX 370 / AMD Radeon 890M (16 CUs, RDNA 3.5, 32GB LPDDR5X UMA).
Test Asset: Multi-column Thai Police Rank Standard Classification Document (test_page2.png).

Backend / Engine Vision Projector (mmproj) Vision Ingestion (Prompt) Generation Speed (Decode) Avg Full Page Time Thai Character Fidelity
๐Ÿš€ Native ROCm / HIP 7.2 (With PR #112) ๐Ÿ›ก๏ธ GPU Accelerated ๐Ÿ† 327.8 tok/s ๐Ÿ† 46.7 tok/s โฑ๏ธ ~12.8 s ๐ŸŽฏ 100% Perfect Match
๐Ÿ”ด Vulkan (SPIR-V Shaders) ๐Ÿ›ก๏ธ GPU Accelerated โšก ~110โ€“130 tok/s ๐Ÿš€ 28.6 tok/s โฑ๏ธ ~29.1 s ๐ŸŽฏ 99.61% Match
โณ Legacy ROCm / HIP (--no-mmproj-offload) ๐Ÿข CPU (AVX-512) ~40โ€“50 tok/s ๐Ÿš€ 45 tok/s โฑ๏ธ ~50.6 s ๐ŸŽฏ 99.61% Match
โŒ Unpatched ROCm / HIP (Without PR #112) GPU (Buggy pointer) Corrupted N/A Invalid โš ๏ธ Severe Hallucinations

๐Ÿ› What was the Bug and Why is PR #112 Required?

In earlier releases of CUDA/HIP backends in llama.cpp and ROCmFPX, running multimodal vision models (such as Qwen2-VL, SigLIP, and Typhoon-OCR) on GPU caused output corruption and hallucinations.

The Root Cause:

In ggml-cuda.cu, intermediate activation buffer src1_ddf_i is allocated as 32-bit float (float *). In ggml_cuda_op_mul_mat_cublas, when multiplying matrices where src1->type == GGML_TYPE_F16, the code mistakenly bypassed the to_fp16_cuda conversion kernel and directly cast the pointer:

// Bug in upstream: Reading 32-bit floats as pairs of 16-bit halfs!
const half * src1_as_half = (const half *) src1_ddf_i;

This caused the GPU GEMM kernel to interpret 32-bit IEEE floats as pairs of 16-bit half floats, corrupting the first Vision Transformer embedding layer by over 40x magnitude.

The Fix (PR #112):

PR #112 enforces explicit to_fp16_cuda type conversion into an allocated half-precision scratch buffer before invoking hipblasGemmEx / cublasGemmEx. This restores 100% mathematical accuracy on the GPU for all vision and multimodal models.


๐Ÿ“ Repository Files

File Size Description
typhoon-ocr1.5-2b-Q4_0_ROCMFP4-imatrix.gguf 1.08 GB Quantized text model backbone (Q4_0_ROCMFP4 with Q6_K token embeddings)
typhoon-ocr1.5-2b.mmproj-f16.gguf 781 MB Full FP16 multimodal vision projector (Required for image/PDF input)

๐Ÿ› ๏ธ Building ROCmFPX with PR #112 (Native HIP)

To get full GPU acceleration on AMD Radeon iGPUs/dGPUs under Windows 11 or Linux:

1. Clone & Apply PR #112:

git clone https://github.com/charlie12345/ROCmFPX.git
cd ROCmFPX

# Fetch and checkout PR #112
git fetch origin pull/112/head:pr-112
git checkout pr-112

# Alternatively, clone the author's fork directly:
# git clone -b fix-cublas-f16-pointer-mismatch https://github.com/nanashi66/ROCmFPX.git

2. Configure & Compile:

  • Windows (ROCm 7.2 / Visual Studio Build Tools + Ninja):

    $env:PATH = "C:\Program Files\AMD\ROCm\7.2\bin;" + $env:PATH
    
    cmake -B build-hip -G Ninja `
      -DGGML_HIP=ON `
      -DAMDGPU_TARGETS="gfx1150;gfx1100" `
      -DCMAKE_BUILD_TYPE=Release `
      -DLLAMA_BUILD_WEBUI=OFF
    
    ninja -C build-hip bin/llama-server.exe bin/llama-cli.exe
    
  • Linux (ROCm 6.x / 7.x):

    cmake -B build-hip -G Ninja \
      -DGGML_HIP=ON \
      -DAMDGPU_TARGETS="gfx1150;gfx1100" \
      -DCMAKE_BUILD_TYPE=Release \
      -DLLAMA_BUILD_WEBUI=OFF
    
    ninja -C build-hip bin/llama-server bin/llama-cli
    

๐Ÿš€ Running Typhoon-OCR 1.5 with Native GPU Acceleration

1. Launch llama-server (Production Web API)

Set the required environment variables and launch the server:

# Windows PowerShell
$env:HSA_OVERRIDE_GFX_VERSION = "11.5.0"  # Required for Strix Point (Radeon 890M / gfx1150)
$env:GGML_CUDA_FORCE_MMQ = "1"

llama-server.exe `
  -m typhoon-ocr1.5-2b-Q4_0_ROCMFP4-imatrix.gguf `
  --mmproj typhoon-ocr1.5-2b.mmproj-f16.gguf `
  -ngl 99 `
  -c 8192 `
  -ctk q8_0 -ctv q8_0 `
  -fa on `
  --port 8083 `
  --host 0.0.0.0

  • -ctk q8_0 -ctv q8_0: Uses 8-bit quantized KV cache to halve VRAM usage with zero degradation.
  • -fa on: Enables Flash Attention to compute long document context in linear memory.
  • Notice: Notice that --no-mmproj-offload is no longer needed!

2. Command-Line OCR (llama-cli)

To test directly from the command line on an image:

llama-cli.exe `
  -m typhoon-ocr1.5-2b-Q4_0_ROCMFP4-imatrix.gguf `
  --mmproj typhoon-ocr1.5-2b.mmproj-f16.gguf `
  --image sample_document.png `
  -ngl 99 `
  -c 8192 `
  -ctk q8_0 -ctv q8_0 `
  -fa on `
  -n 1024 `
  -p "Extract all text from the image.\n\nInstructions:\n- Only return the clean Markdown.\n- Do not include any explanation or extra text.\n- You must include all information on the page.\n\nFormatting Rules:\n- Tables: Render tables using <table>...</table> in clean HTML format.\n- Equations: Render equations using LaTeX syntax with inline ($...$) and block ($$...$$)."

3. Python Integration (Official OpenAI Compatible API)

import base64
import requests

def ocr_document(image_path: str, server_url: str = "http://localhost:8083/v1/chat/completions"):
    with open(image_path, "rb") as f:
        img_b64 = base64.b64encode(f.read()).decode("utf-8")

    prompt = """Extract all text from the image.
Instructions:
- Only return the clean Markdown.
- Do not include any explanation or extra text.
- You must include all information on the page.
Formatting Rules:
- Tables: Render tables using <table>...</table> in clean HTML format.
- Equations: Render equations using LaTeX syntax with inline ($...$) and block ($$...$$).
- Page Numbers: Wrap page numbers in <page_number>...</page_number>."""

    payload = {
        "model": "typhoon-ocr1.5-2b",
        "messages": [
            {
                "role": "user",
                "content": [
                    {"type": "text", "text": prompt},
                    {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img_b64}"}}
                ]
            }
        ],
        "temperature": 0.0,
        "max_tokens": 4096
    }

    response = requests.post(server_url, json=payload, timeout=120)
    return response.json()["choices"][0]["message"]["content"]

# Example:
# print(ocr_document("thai_legal_doc.png"))

๐Ÿ’– Credits & Acknowledgements

  1. Original Model: SCB 10X / Typhoon AI Team for typhoon-ai/typhoon-ocr1.5-2b.
  2. Upstream ROCm Engine & Fork: charlie12345 for the ROCmFPX engine.
  3. Bug Fix & Multimodal Enablement: nanashi66 via ROCmFPX PR #112 for discovering and fixing the Vision Transformer cuBLAS/hipBLAS FP16 pointer mismatch.
  4. Foundational Framework: Georgi Gerganov & the llama.cpp open-source community.

โ˜• Support Our Work

If you find this ROCmFP4 quantization, empirical benchmark data, and Native ROCm/HIP GPU bugfixes helpful for your local AI workflows on AMD hardware, consider supporting our ongoing research and optimization efforts:

  • Bitcoin (BTC): bc1pm9yqcj5m5jv3jzxkhv4uvrgl2weyq46m2jlzufzyfwyr4fmv8gwq548kw6

Your support helps us continue optimizing open-weight models, submitting upstream ROCm/HIP patches, and developing high-efficiency inference recipes for AMD Strix Point and RDNA architectures!


๐Ÿ“œ License

Apache 2.0

Downloads last month
320
GGUF
Model size
2B params
Architecture
qwen3vl
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF

Quantized
(6)
this model