Instructions to use nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: llama cli -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: llama cli -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: ./llama-cli -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: ./build/bin/llama-cli -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
Use Docker
docker model run hf.co/nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
- LM Studio
- Jan
- vLLM
How to use nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
- Ollama
How to use nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF with Ollama:
ollama run hf.co/nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
- Unsloth Desktop
- Pi
How to use nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF with Docker Model Runner:
docker model run hf.co/nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
- Lemonade
How to use nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
Run and chat with the model
lemonade run user.typhoon-ocr1.5-2b-ROCMFP4-GGUF-Q4_0_ROCMFP
List all available models
lemonade list
- Hermes Agent
How to use nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure the model in Pi
# Install Pi:
npm install -g @earendil-works/pi-coding-agent# Add to ~/.pi/agent/models.json:
{
"providers": {
"llama-cpp": {
"baseUrl": "http://localhost:8080/v1",
"api": "openai-completions",
"apiKey": "none",
"models": [
{
"id": "nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP"
}
]
}
}
}Run Pi
# Start Pi in your project directory:
pi- Typhoon-OCR 1.5 2B โ ROCmFP4 GGUF (Thai iMatrix Calibrated)
- ๐ Key Highlights
- ๐ Benchmark & Quality Verification (Measured on AMD Radeon 890M / gfx1150)
- ๐ What was the Bug and Why is PR #112 Required?
- ๐ Repository Files
- ๐ ๏ธ Building ROCmFPX with PR #112 (Native HIP)
- ๐ Running Typhoon-OCR 1.5 with Native GPU Acceleration
- ๐ Credits & Acknowledgements
- โ Support Our Work
- ๐ License
- ๐ Key Highlights
Typhoon-OCR 1.5 2B โ ROCmFP4 GGUF (Thai iMatrix Calibrated)
This repository provides ROCmFP4 quantized GGUF weights for typhoon-ai/typhoon-ocr1.5-2b, calibrated specifically for high-fidelity Thai document OCR using Importance Matrix (iMatrix).
The text backbone is quantized to Q4_0_ROCMFP4 (UE4M3-scale experimental with Q6_K token embeddings), while the Vision Projector (mmproj-f16.gguf) remains unquantized in full FP16 to ensure zero loss in visual document resolution.
๐ UPDATE (Sep 2026): 100% Native ROCm / HIP GPU Acceleration is LIVE!
Full GPU acceleration (Vision Transformer + LLM Backbone) is now supported natively via PR #112 (charlie12345/ROCmFPX#112)!
- No More CPU Fallback: You no longer need
--no-mmproj-offload! Both the vision encoder and LLM run 100% on the AMD GPU.- Blazing Speeds on AMD Radeon 890M (Strix Point):
- โก Vision Prompt Processing:
327.8 tok/s(entire high-res document ingested in ~6โ7 seconds)- ๐ Text Generation (OCR Decode):
46.7 tok/s- ๐พ Ultra-Low VRAM: Total memory footprint is only ~2.7 GB (allowing large batches or multi-turn OCR on 16GBโ32GB unified memory laptops).
- Build Requirement: Until PR #112 is merged into upstream master releases, you must build the engine from source using PR #112 or the branch
nanashi66:fix-cublas-f16-pointer-mismatch. See the Build Instructions below.- Alternative (Plug-and-Play): If you prefer not to build from source, you can still use the pre-built Vulkan backend of
ROCmFPXwhich runs out of the box.
๐ Key Highlights
- 41% Smaller than Q8_0 (67% Smaller than FP16): Text backbone compressed down to 1.08 GB (from 3.28 GB FP16).
- High-Precision Thai Preservation: Calibrated on a custom Thai Calibration Mix (70% general Thai corpus + 30% complex Thai official & legal documents), preventing degradation of Thai vowels, tone marks (เธงเธฃเธฃเธเธขเธธเธเธเน), and specialized vocabulary.
- 100% Empirical Document Fidelity: Verified on complex Thai government and legal documents with exact table preservation (
<table>...</table>), zero character hallucinations, and 100% identical Thai numerals. - Tuned for AMD RDNA 3 / 3.5 (
gfx1150/gfx1100/gfx1103): High inference speed (46.7 tok/s) and ultra-low VRAM footprint (~2.7 GB) on AMD Ryzen AI 9 HX 370 / Radeon 890M / 880M / 780M iGPUs.
๐ Benchmark & Quality Verification (Measured on AMD Radeon 890M / gfx1150)
Hardware: AMD Ryzen AI 9 HX 370 / AMD Radeon 890M (16 CUs, RDNA 3.5, 32GB LPDDR5X UMA).
Test Asset: Multi-column Thai Police Rank Standard Classification Document (test_page2.png).
| Backend / Engine | Vision Projector (mmproj) |
Vision Ingestion (Prompt) | Generation Speed (Decode) | Avg Full Page Time | Thai Character Fidelity |
|---|---|---|---|---|---|
| ๐ Native ROCm / HIP 7.2 (With PR #112) | ๐ก๏ธ GPU Accelerated | ๐ 327.8 tok/s |
๐ 46.7 tok/s |
โฑ๏ธ ~12.8 s | ๐ฏ 100% Perfect Match |
| ๐ด Vulkan (SPIR-V Shaders) | ๐ก๏ธ GPU Accelerated | โก ~110โ130 tok/s | ๐ 28.6 tok/s |
โฑ๏ธ ~29.1 s | ๐ฏ 99.61% Match |
โณ Legacy ROCm / HIP (--no-mmproj-offload) |
๐ข CPU (AVX-512) | ~40โ50 tok/s | ๐ 45 tok/s |
โฑ๏ธ ~50.6 s | ๐ฏ 99.61% Match |
| โ Unpatched ROCm / HIP (Without PR #112) | GPU (Buggy pointer) | Corrupted | N/A | Invalid | โ ๏ธ Severe Hallucinations |
๐ What was the Bug and Why is PR #112 Required?
In earlier releases of CUDA/HIP backends in llama.cpp and ROCmFPX, running multimodal vision models (such as Qwen2-VL, SigLIP, and Typhoon-OCR) on GPU caused output corruption and hallucinations.
The Root Cause:
In ggml-cuda.cu, intermediate activation buffer src1_ddf_i is allocated as 32-bit float (float *). In ggml_cuda_op_mul_mat_cublas, when multiplying matrices where src1->type == GGML_TYPE_F16, the code mistakenly bypassed the to_fp16_cuda conversion kernel and directly cast the pointer:
// Bug in upstream: Reading 32-bit floats as pairs of 16-bit halfs!
const half * src1_as_half = (const half *) src1_ddf_i;
This caused the GPU GEMM kernel to interpret 32-bit IEEE floats as pairs of 16-bit half floats, corrupting the first Vision Transformer embedding layer by over 40x magnitude.
The Fix (PR #112):
PR #112 enforces explicit to_fp16_cuda type conversion into an allocated half-precision scratch buffer before invoking hipblasGemmEx / cublasGemmEx. This restores 100% mathematical accuracy on the GPU for all vision and multimodal models.
๐ Repository Files
| File | Size | Description |
|---|---|---|
typhoon-ocr1.5-2b-Q4_0_ROCMFP4-imatrix.gguf |
1.08 GB | Quantized text model backbone (Q4_0_ROCMFP4 with Q6_K token embeddings) |
typhoon-ocr1.5-2b.mmproj-f16.gguf |
781 MB | Full FP16 multimodal vision projector (Required for image/PDF input) |
๐ ๏ธ Building ROCmFPX with PR #112 (Native HIP)
To get full GPU acceleration on AMD Radeon iGPUs/dGPUs under Windows 11 or Linux:
1. Clone & Apply PR #112:
git clone https://github.com/charlie12345/ROCmFPX.git
cd ROCmFPX
# Fetch and checkout PR #112
git fetch origin pull/112/head:pr-112
git checkout pr-112
# Alternatively, clone the author's fork directly:
# git clone -b fix-cublas-f16-pointer-mismatch https://github.com/nanashi66/ROCmFPX.git
2. Configure & Compile:
Windows (ROCm 7.2 / Visual Studio Build Tools + Ninja):
$env:PATH = "C:\Program Files\AMD\ROCm\7.2\bin;" + $env:PATH cmake -B build-hip -G Ninja ` -DGGML_HIP=ON ` -DAMDGPU_TARGETS="gfx1150;gfx1100" ` -DCMAKE_BUILD_TYPE=Release ` -DLLAMA_BUILD_WEBUI=OFF ninja -C build-hip bin/llama-server.exe bin/llama-cli.exeLinux (ROCm 6.x / 7.x):
cmake -B build-hip -G Ninja \ -DGGML_HIP=ON \ -DAMDGPU_TARGETS="gfx1150;gfx1100" \ -DCMAKE_BUILD_TYPE=Release \ -DLLAMA_BUILD_WEBUI=OFF ninja -C build-hip bin/llama-server bin/llama-cli
๐ Running Typhoon-OCR 1.5 with Native GPU Acceleration
1. Launch llama-server (Production Web API)
Set the required environment variables and launch the server:
# Windows PowerShell
$env:HSA_OVERRIDE_GFX_VERSION = "11.5.0" # Required for Strix Point (Radeon 890M / gfx1150)
$env:GGML_CUDA_FORCE_MMQ = "1"
llama-server.exe `
-m typhoon-ocr1.5-2b-Q4_0_ROCMFP4-imatrix.gguf `
--mmproj typhoon-ocr1.5-2b.mmproj-f16.gguf `
-ngl 99 `
-c 8192 `
-ctk q8_0 -ctv q8_0 `
-fa on `
--port 8083 `
--host 0.0.0.0
-ctk q8_0 -ctv q8_0: Uses 8-bit quantized KV cache to halve VRAM usage with zero degradation.-fa on: Enables Flash Attention to compute long document context in linear memory.- Notice: Notice that
--no-mmproj-offloadis no longer needed!
2. Command-Line OCR (llama-cli)
To test directly from the command line on an image:
llama-cli.exe `
-m typhoon-ocr1.5-2b-Q4_0_ROCMFP4-imatrix.gguf `
--mmproj typhoon-ocr1.5-2b.mmproj-f16.gguf `
--image sample_document.png `
-ngl 99 `
-c 8192 `
-ctk q8_0 -ctv q8_0 `
-fa on `
-n 1024 `
-p "Extract all text from the image.\n\nInstructions:\n- Only return the clean Markdown.\n- Do not include any explanation or extra text.\n- You must include all information on the page.\n\nFormatting Rules:\n- Tables: Render tables using <table>...</table> in clean HTML format.\n- Equations: Render equations using LaTeX syntax with inline ($...$) and block ($$...$$)."
3. Python Integration (Official OpenAI Compatible API)
import base64
import requests
def ocr_document(image_path: str, server_url: str = "http://localhost:8083/v1/chat/completions"):
with open(image_path, "rb") as f:
img_b64 = base64.b64encode(f.read()).decode("utf-8")
prompt = """Extract all text from the image.
Instructions:
- Only return the clean Markdown.
- Do not include any explanation or extra text.
- You must include all information on the page.
Formatting Rules:
- Tables: Render tables using <table>...</table> in clean HTML format.
- Equations: Render equations using LaTeX syntax with inline ($...$) and block ($$...$$).
- Page Numbers: Wrap page numbers in <page_number>...</page_number>."""
payload = {
"model": "typhoon-ocr1.5-2b",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img_b64}"}}
]
}
],
"temperature": 0.0,
"max_tokens": 4096
}
response = requests.post(server_url, json=payload, timeout=120)
return response.json()["choices"][0]["message"]["content"]
# Example:
# print(ocr_document("thai_legal_doc.png"))
๐ Credits & Acknowledgements
- Original Model: SCB 10X / Typhoon AI Team for
typhoon-ai/typhoon-ocr1.5-2b. - Upstream ROCm Engine & Fork: charlie12345 for the ROCmFPX engine.
- Bug Fix & Multimodal Enablement: nanashi66 via ROCmFPX PR #112 for discovering and fixing the Vision Transformer cuBLAS/hipBLAS FP16 pointer mismatch.
- Foundational Framework: Georgi Gerganov & the
llama.cppopen-source community.
โ Support Our Work
If you find this ROCmFP4 quantization, empirical benchmark data, and Native ROCm/HIP GPU bugfixes helpful for your local AI workflows on AMD hardware, consider supporting our ongoing research and optimization efforts:
- Bitcoin (BTC):
bc1pm9yqcj5m5jv3jzxkhv4uvrgl2weyq46m2jlzufzyfwyr4fmv8gwq548kw6
Your support helps us continue optimizing open-weight models, submitting upstream ROCm/HIP patches, and developing high-efficiency inference recipes for AMD Strix Point and RDNA architectures!
๐ License
Apache 2.0
- Downloads last month
- 320
4-bit
Model tree for nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF
Base model
Qwen/Qwen3-VL-2B-Instruct
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp# Start a local OpenAI-compatible server: llama serve -hf nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF:Q4_0_ROCMFP