CHAI OCR Logo

CHAI OCR

Qwen 3.5 4B VL · Vision-Boost α=1.25 · Task-Arithmetic Merge

Experimental OCR and Document Understanding Model


Overview

CHAI OCR is an experimental vision-language model derived from Qwen 3.5 4B VL and optimized toward Optical Character Recognition (OCR) and document understanding.

The model uses a task-arithmetic merge that selectively amplifies the vision-tower adaptation of a vision-tuned Qwen 3.5 checkpoint while preserving the language-model components of the tuned source.

Evaluation status: Experimental / unvalidated.

The current configuration represents a hypothesis that must be evaluated against the original Qwen 3.5 4B VL checkpoint under identical benchmark and decoding conditions.

The unmodified Qwen 3.5 4B VL checkpoint should therefore be treated as the control model in all comparisons.


Model Architecture

CHAI OCR applies task-vector extrapolation to the vision tower:

W_out = W_qwen + α × (W_tuned − W_qwen)

The merge coefficients are:

model.visual.*     α = 1.25
everything else    α = 1.00

This means:

Component Alpha Modification
Vision Tower 1.25 Vision adaptation amplified by 25%
Language Model 1.00 Identical to vision-tuned source
MTP Head 1.00 Identical to vision-tuned source
Embeddings 1.00 Identical to vision-tuned source

Only the 297 vision-tower tensors, representing approximately 333 million parameters, are modified relative to the vision-tuned source.

The language model, MTP head, and embeddings remain bit-identical to the tuned source, subject to the precision-handling details described below.


Benchmark Results

Current CHAI OCR benchmark results:

Model Score
CHAI OCR (final) 78.6
DeepSeek-OCR-2 76.3
LightOnOCR-1B-1025 76.1
DeepSeek-OCR 75.7
MinerU2.5-2509-1.2B 75.2
GLM-OCR 75.2
kraken-ppocrv6-medium 36.7741

Benchmark scores should be interpreted only within the corresponding evaluation configuration. Reproduction should use identical prompts, preprocessing, decoding parameters, seeds, and benchmark versions.


Why Modify the Vision Tower?

The design of CHAI OCR is motivated by a tensor-level comparison between the vision-tuned source checkpoint and the original Qwen 3.5 model.

Two observations were particularly important.

1. Limited Vision-Tower Adaptation

The vision tower showed significantly less parameter movement during fine-tuning than the language model:

Vision tower relative drift:     1.49%
Language model relative drift:   9.08%
Vision LayerNorm drift:          0.04%

The extremely small LayerNorm movement suggests that substantial portions of the visual representation stack remained close to the original Qwen checkpoint.

2. Weak Performance on Degraded Scans

The source model's weakest OCR benchmark category was:

old_scans:      51.1
multi_column:   82.1
baseline:       99.9

The old_scans result was approximately 31 points below the next-worst evaluated category.

Degraded-document recognition is heavily dependent on the quality and robustness of visual representations.

This creates the working hypothesis behind CHAI OCR:

The weakest capability is strongly vision-dependent, while the vision component also appears to have received the least adaptation during training.

CHAI OCR therefore moves the vision parameters an additional 25% along the task vector already learned during fine-tuning.

Mathematically:

Δvision = W_tuned − W_qwen

W_CHAI = W_qwen + 1.25 × Δvision

This is an extrapolation rather than additional training.

Important Limitation

The relationship above is a correlation, not proof of causation.

Task-vector extrapolation may improve the desired capability, leave it unchanged, or degrade unrelated model capabilities.

Controlled evaluation is therefore required before treating the modification as an improvement.


Model Provenance

Property Value
Model Name CHAI OCR
Architecture Qwen3.5 Vision-Language Model
Base Architecture Qwen3_5ForConditionalGeneration
Base Checkpoint MergeKit/Qwen / Qwen 3.5
Fine-Tuned Source Vision-tuned Qwen 3.5 checkpoint
Merge Method Task-vector arithmetic
Vision Alpha 1.25
Non-Vision Alpha 1.00
Merge Tool MergeKit/merge_task_vector.py
Merge Recipe merge_recipe.json
Primary Task OCR / Document Understanding
Framework Hugging Face Transformers

The merge_recipe.json file records the per-tensor configuration, including:

  • tensor names;
  • alpha values;
  • tensor shapes;
  • data types; and
  • merge configuration.

Merge Validation

The merge implementation was tested for reproducibility at known interpolation points.

α = 1

W_out = W_qwen + 1 × (W_tuned − W_qwen)
      = W_tuned

The tool reproduces the tuned source bit-exactly.

Validation:

357 tensors verified

α = 0

W_out = W_qwen

The merge reproduces the original Qwen parameters.

Fractional Alpha

Fractional task-vector coefficients were independently recomputed and compared across shard boundaries.

Validation:

429 tensors verified

The resulting tensors matched the independent recomputation bit-exactly.


Precision Handling

CHAI OCR intentionally introduces two implementation differences from the source checkpoint files.

Float32 SSM Tensors

The 48 SSM tensors are stored as float32 rather than bfloat16.

The vision-tuned source had downcast:

linear_attn.A_log
linear_attn.norm.weight

from Qwen's original float32 representation.

However, the model configuration declares:

"mamba_ssm_dtype": "float32"

CHAI OCR therefore restores these tensors to the base model's intended precision instead of propagating the source checkpoint's BF16 downcast.

Tied Language-Model Head

lm_head.weight is stored as a copy of the merged embedding matrix.

The configuration specifies:

tie_word_embeddings = true

The source checkpoint's lm_head is also bit-identical to its corresponding embed_tokens tensor.

CHAI OCR therefore explicitly preserves this weight relationship.


Hardware Requirements

The model contains approximately:

~10.6 GB of BF16 model weights

Recommended GPU memory:

GPU VRAM Expected Use
12 GB Very constrained
16 GB Workable
24 GB Recommended
32 GB+ Comfortable for larger inference workloads

Actual VRAM consumption depends on:

  • input resolution;
  • number of pages or images;
  • attention implementation;
  • generation length;
  • KV-cache size;
  • batching;
  • inference framework; and
  • quantization.

The model was assembled on hardware that does not have sufficient VRAM to run the complete BF16 checkpoint directly.


Installation

Install the primary dependencies:

pip install -U transformers accelerate

Additional dependencies may be required depending on the inference pipeline.

For production inference with vLLM:

pip install -U vllm

Running CHAI OCR

The following assets are inherited from the vision-tuned source:

  • processor;
  • tokenizer;
  • chat template;
  • model configuration; and
  • multimodal configuration.

The model therefore loads using the standard Qwen 3.5 vision-language workflow.

Transformers

python your_vl_runner.py input.pdf \
    --model_path /path/to/chai-ocr

Example Python initialization:

from transformers import AutoProcessor, AutoModelForImageTextToText

model_path = "/path/to/chai-ocr"

processor = AutoProcessor.from_pretrained(
    model_path,
    trust_remote_code=True
)

model = AutoModelForImageTextToText.from_pretrained(
    model_path,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True
)

vLLM Deployment

CHAI OCR can also be exposed as an inference service using vLLM:

vllm serve /path/to/chai-ocr \
    --served-model-name chai-ocr

The deployed model can then be accessed through the compatible inference API supported by the installed vLLM version.


Recommended Evaluation

CHAI OCR should be benchmarked against the unmodified Qwen 3.5 4B VL checkpoint using identical evaluation conditions.

A useful evaluation suite is:

allenai/olmOCR-bench

Important controls include:

same benchmark revision
same image preprocessing
same prompts
same system prompt
same decoding settings
same temperature
same maximum output length
same random seed

Priority Metrics

Task Reference Baseline Purpose
old_scans 51.1 Primary target
overall 85.8 Global regression check
baseline 99.9 High-sensitivity canary
arxiv_math 86.9 Language/structured-content canary

The principal experimental question is:

Does CHAI OCR improve old_scans performance without materially reducing overall, baseline, or other high-performing capabilities?

An improvement in degraded scans that is offset by a substantial regression elsewhere should not be interpreted as a successful model improvement.


Evaluation Strategy

A recommended evaluation order is:

1. baseline
2. overall
3. old_scans
4. arxiv_math
5. multi_column
6. complete benchmark suite

The baseline category acts as an early canary.

If the score falls materially from the reference value of 99.9, particularly below approximately 99, the task-vector extrapolation may have damaged capabilities that were already close to saturation.

In such a case, broader evaluation should be interpreted cautiously.


Alpha Sweep

If α=1.25 provides a measurable improvement, the next step is to evaluate several task-vector strengths.

For example:

python merge_task_vector.py \
    --alpha "vision=1.10" \
    --out ./chai-ocr-vb-1.10

python merge_task_vector.py \
    --alpha "vision=1.15" \
    --out ./chai-ocr-vb-1.15

python merge_task_vector.py \
    --alpha "vision=1.25" \
    --out ./chai-ocr-vb-1.25

python merge_task_vector.py \
    --alpha "vision=1.35" \
    --out ./chai-ocr-vb-1.35

python merge_task_vector.py \
    --alpha "vision=1.50" \
    --out ./chai-ocr-vb-1.50

Results should be plotted against α for both the target task and regression-sensitive tasks.

Conceptually:

                  old_scans ↑
                      │
                      │        ●
                      │     ●
                      │  ●
                      │
                      └────────────────→ α
                         1.0  1.25  1.5

The optimal coefficient is not necessarily the largest one.

Task-vector extrapolation generally becomes increasingly unstable as parameters are moved farther outside the trained checkpoint.

Values around or above α ≈ 1.5 should therefore be treated as exploratory rather than assumed improvements.

If performance continues increasing at that point, the more defensible conclusion may be that the vision stack requires additional targeted training rather than stronger weight extrapolation.


Improving Degraded Scan Recognition

Task arithmetic is only one possible approach to improving old_scans.

Image preprocessing may provide a more predictable improvement for degraded documents.

Potential preprocessing stages include:

Input Scan
    │
    ▼
Orientation Detection
    │
    ▼
Deskew
    │
    ▼
Denoising
    │
    ▼
Local Contrast Enhancement
    │
    ▼
Adaptive Binarization
    │
    ▼
Resolution Normalization
    │
    ▼
CHAI OCR

Useful preprocessing techniques include:

  • deskewing;
  • denoising;
  • CLAHE contrast enhancement;
  • adaptive thresholding;
  • Sauvola binarization;
  • background normalization;
  • border removal;
  • orientation correction; and
  • intelligent resizing.

For heavily degraded scans, preprocessing should be benchmarked independently as well as in combination with CHAI OCR.


Intended Use

CHAI OCR is intended for research and experimentation involving:

  • Optical Character Recognition;
  • scanned-document transcription;
  • document understanding;
  • historical document processing;
  • degraded document recognition;
  • document layout interpretation;
  • tables and forms;
  • multi-column documents;
  • visually structured text;
  • document digitization; and
  • multimodal document AI research.

Experimental Status

CHAI OCR should currently be considered an experimental research checkpoint.

The task-arithmetic modification has a plausible technical motivation, but the effect of α=1.25 must be established through controlled evaluation.

Users should therefore avoid assuming that:

larger α = better OCR

or that improvements on one benchmark category necessarily generalize to all document types.


Reproducibility

For reproducible comparisons, record at minimum:

Model revision
Benchmark revision
Transformers version
PyTorch version
CUDA version
GPU model
Image resolution
Image preprocessing
Prompt template
Generation parameters
Random seed
Batch size
Attention backend

Benchmark comparisons should not mix materially different preprocessing or decoding configurations.


Project Structure

A typical CHAI OCR repository may use the following structure:

chai-ocr/
│
├── README.md
├── config.json
├── generation_config.json
├── tokenizer.json
├── tokenizer_config.json
├── processor_config.json
├── preprocessor_config.json
├── chat_template.json
│
├── model-00001-of-000xx.safetensors
├── model-00002-of-000xx.safetensors
├── ...
├── model.safetensors.index.json
│
├── chai-ocr.png
│
├── merge_recipe.json
├── merge_task_vector.py
│
├── benchmark/
│   ├── olmocr/
│   ├── results/
│   └── evaluation_config.json
│
└── examples/
    ├── images/
    ├── documents/
    └── inference.py

Summary

CHAI OCR explores a narrowly targeted modification to a Qwen 3.5 vision-language OCR checkpoint:

Qwen 3.5 Base
      │
      ├───────────────┐
      │               │
      ▼               ▼
Vision Tower      Language Model
      │               │
      │             α = 1.0
      │               │
    α = 1.25           │
      │               │
      └───────┬───────┘
              ▼
           CHAI OCR

The experiment is based on three observations:

  1. The vision tower changed substantially less than the language model during the source fine-tuning.
  2. Degraded scans were the source checkpoint's weakest measured OCR category.
  3. The vision tower is the component most directly responsible for extracting useful representations from degraded page images.

CHAI OCR therefore tests whether modest extrapolation of the learned vision task vector can improve degraded-document OCR while preserving the capabilities of the tuned language model.

The central criterion remains:

Improve difficult visual OCR without sacrificing the capabilities that already work well.


CHAI OCR

Vision-Language OCR · Document Intelligence · Experimental Research

Downloads last month
218
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support