Instructions to use surendirakrishna/salad with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use surendirakrishna/salad with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="surendirakrishna/salad") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("surendirakrishna/salad") model = AutoModelForMultimodalLM.from_pretrained("surendirakrishna/salad", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use surendirakrishna/salad with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "surendirakrishna/salad" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "surendirakrishna/salad", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/surendirakrishna/salad
- SGLang
How to use surendirakrishna/salad with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "surendirakrishna/salad" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "surendirakrishna/salad", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "surendirakrishna/salad" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "surendirakrishna/salad", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use surendirakrishna/salad with Docker Model Runner:
docker model run hf.co/surendirakrishna/salad
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM
processor = AutoProcessor.from_pretrained("surendirakrishna/salad")
model = AutoModelForMultimodalLM.from_pretrained("surendirakrishna/salad", device_map="auto")
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
{"type": "text", "text": "What animal is on the candy?"}
]
},
]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:]))- CHAI OCR
- Overview
- Model Architecture
- Benchmark Results
- Why Modify the Vision Tower?
- Model Provenance
- Merge Validation
- Precision Handling
- Hardware Requirements
- Installation
- Running CHAI OCR
- vLLM Deployment
- Recommended Evaluation
- Evaluation Strategy
- Alpha Sweep
- Improving Degraded Scan Recognition
- Intended Use
- Experimental Status
- Reproducibility
- Project Structure
- Summary
- Overview
CHAI OCR
Qwen 3.5 4B VL · Vision-Boost α=1.25 · Task-Arithmetic Merge
Experimental OCR and Document Understanding Model
Overview
CHAI OCR is an experimental vision-language model derived from Qwen 3.5 4B VL and optimized toward Optical Character Recognition (OCR) and document understanding.
The model uses a task-arithmetic merge that selectively amplifies the vision-tower adaptation of a vision-tuned Qwen 3.5 checkpoint while preserving the language-model components of the tuned source.
Evaluation status: Experimental / unvalidated.
The current configuration represents a hypothesis that must be evaluated against the original Qwen 3.5 4B VL checkpoint under identical benchmark and decoding conditions.
The unmodified Qwen 3.5 4B VL checkpoint should therefore be treated as the control model in all comparisons.
Model Architecture
CHAI OCR applies task-vector extrapolation to the vision tower:
W_out = W_qwen + α × (W_tuned − W_qwen)
The merge coefficients are:
model.visual.* α = 1.25
everything else α = 1.00
This means:
| Component | Alpha | Modification |
|---|---|---|
| Vision Tower | 1.25 | Vision adaptation amplified by 25% |
| Language Model | 1.00 | Identical to vision-tuned source |
| MTP Head | 1.00 | Identical to vision-tuned source |
| Embeddings | 1.00 | Identical to vision-tuned source |
Only the 297 vision-tower tensors, representing approximately 333 million parameters, are modified relative to the vision-tuned source.
The language model, MTP head, and embeddings remain bit-identical to the tuned source, subject to the precision-handling details described below.
Benchmark Results
Current CHAI OCR benchmark results:
| Model | Score |
|---|---|
| CHAI OCR (final) | 78.6 |
| DeepSeek-OCR-2 | 76.3 |
| LightOnOCR-1B-1025 | 76.1 |
| DeepSeek-OCR | 75.7 |
| MinerU2.5-2509-1.2B | 75.2 |
| GLM-OCR | 75.2 |
| kraken-ppocrv6-medium | 36.7741 |
Benchmark scores should be interpreted only within the corresponding evaluation configuration. Reproduction should use identical prompts, preprocessing, decoding parameters, seeds, and benchmark versions.
Why Modify the Vision Tower?
The design of CHAI OCR is motivated by a tensor-level comparison between the vision-tuned source checkpoint and the original Qwen 3.5 model.
Two observations were particularly important.
1. Limited Vision-Tower Adaptation
The vision tower showed significantly less parameter movement during fine-tuning than the language model:
Vision tower relative drift: 1.49%
Language model relative drift: 9.08%
Vision LayerNorm drift: 0.04%
The extremely small LayerNorm movement suggests that substantial portions of the visual representation stack remained close to the original Qwen checkpoint.
2. Weak Performance on Degraded Scans
The source model's weakest OCR benchmark category was:
old_scans: 51.1
multi_column: 82.1
baseline: 99.9
The old_scans result was approximately 31 points below the next-worst evaluated category.
Degraded-document recognition is heavily dependent on the quality and robustness of visual representations.
This creates the working hypothesis behind CHAI OCR:
The weakest capability is strongly vision-dependent, while the vision component also appears to have received the least adaptation during training.
CHAI OCR therefore moves the vision parameters an additional 25% along the task vector already learned during fine-tuning.
Mathematically:
Δvision = W_tuned − W_qwen
W_CHAI = W_qwen + 1.25 × Δvision
This is an extrapolation rather than additional training.
Important Limitation
The relationship above is a correlation, not proof of causation.
Task-vector extrapolation may improve the desired capability, leave it unchanged, or degrade unrelated model capabilities.
Controlled evaluation is therefore required before treating the modification as an improvement.
Model Provenance
| Property | Value |
|---|---|
| Model Name | CHAI OCR |
| Architecture | Qwen3.5 Vision-Language Model |
| Base Architecture | Qwen3_5ForConditionalGeneration |
| Base Checkpoint | MergeKit/Qwen / Qwen 3.5 |
| Fine-Tuned Source | Vision-tuned Qwen 3.5 checkpoint |
| Merge Method | Task-vector arithmetic |
| Vision Alpha | 1.25 |
| Non-Vision Alpha | 1.00 |
| Merge Tool | MergeKit/merge_task_vector.py |
| Merge Recipe | merge_recipe.json |
| Primary Task | OCR / Document Understanding |
| Framework | Hugging Face Transformers |
The merge_recipe.json file records the per-tensor configuration, including:
- tensor names;
- alpha values;
- tensor shapes;
- data types; and
- merge configuration.
Merge Validation
The merge implementation was tested for reproducibility at known interpolation points.
α = 1
W_out = W_qwen + 1 × (W_tuned − W_qwen)
= W_tuned
The tool reproduces the tuned source bit-exactly.
Validation:
357 tensors verified
α = 0
W_out = W_qwen
The merge reproduces the original Qwen parameters.
Fractional Alpha
Fractional task-vector coefficients were independently recomputed and compared across shard boundaries.
Validation:
429 tensors verified
The resulting tensors matched the independent recomputation bit-exactly.
Precision Handling
CHAI OCR intentionally introduces two implementation differences from the source checkpoint files.
Float32 SSM Tensors
The 48 SSM tensors are stored as float32 rather than bfloat16.
The vision-tuned source had downcast:
linear_attn.A_log
linear_attn.norm.weight
from Qwen's original float32 representation.
However, the model configuration declares:
"mamba_ssm_dtype": "float32"
CHAI OCR therefore restores these tensors to the base model's intended precision instead of propagating the source checkpoint's BF16 downcast.
Tied Language-Model Head
lm_head.weight is stored as a copy of the merged embedding matrix.
The configuration specifies:
tie_word_embeddings = true
The source checkpoint's lm_head is also bit-identical to its corresponding embed_tokens tensor.
CHAI OCR therefore explicitly preserves this weight relationship.
Hardware Requirements
The model contains approximately:
~10.6 GB of BF16 model weights
Recommended GPU memory:
| GPU VRAM | Expected Use |
|---|---|
| 12 GB | Very constrained |
| 16 GB | Workable |
| 24 GB | Recommended |
| 32 GB+ | Comfortable for larger inference workloads |
Actual VRAM consumption depends on:
- input resolution;
- number of pages or images;
- attention implementation;
- generation length;
- KV-cache size;
- batching;
- inference framework; and
- quantization.
The model was assembled on hardware that does not have sufficient VRAM to run the complete BF16 checkpoint directly.
Installation
Install the primary dependencies:
pip install -U transformers accelerate
Additional dependencies may be required depending on the inference pipeline.
For production inference with vLLM:
pip install -U vllm
Running CHAI OCR
The following assets are inherited from the vision-tuned source:
- processor;
- tokenizer;
- chat template;
- model configuration; and
- multimodal configuration.
The model therefore loads using the standard Qwen 3.5 vision-language workflow.
Transformers
python your_vl_runner.py input.pdf \
--model_path /path/to/chai-ocr
Example Python initialization:
from transformers import AutoProcessor, AutoModelForImageTextToText
model_path = "/path/to/chai-ocr"
processor = AutoProcessor.from_pretrained(
model_path,
trust_remote_code=True
)
model = AutoModelForImageTextToText.from_pretrained(
model_path,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True
)
vLLM Deployment
CHAI OCR can also be exposed as an inference service using vLLM:
vllm serve /path/to/chai-ocr \
--served-model-name chai-ocr
The deployed model can then be accessed through the compatible inference API supported by the installed vLLM version.
Recommended Evaluation
CHAI OCR should be benchmarked against the unmodified Qwen 3.5 4B VL checkpoint using identical evaluation conditions.
A useful evaluation suite is:
Important controls include:
same benchmark revision
same image preprocessing
same prompts
same system prompt
same decoding settings
same temperature
same maximum output length
same random seed
Priority Metrics
| Task | Reference Baseline | Purpose |
|---|---|---|
old_scans |
51.1 | Primary target |
overall |
85.8 | Global regression check |
baseline |
99.9 | High-sensitivity canary |
arxiv_math |
86.9 | Language/structured-content canary |
The principal experimental question is:
Does CHAI OCR improve
old_scansperformance without materially reducingoverall,baseline, or other high-performing capabilities?
An improvement in degraded scans that is offset by a substantial regression elsewhere should not be interpreted as a successful model improvement.
Evaluation Strategy
A recommended evaluation order is:
1. baseline
2. overall
3. old_scans
4. arxiv_math
5. multi_column
6. complete benchmark suite
The baseline category acts as an early canary.
If the score falls materially from the reference value of 99.9, particularly below approximately 99, the task-vector extrapolation may have damaged capabilities that were already close to saturation.
In such a case, broader evaluation should be interpreted cautiously.
Alpha Sweep
If α=1.25 provides a measurable improvement, the next step is to evaluate several task-vector strengths.
For example:
python merge_task_vector.py \
--alpha "vision=1.10" \
--out ./chai-ocr-vb-1.10
python merge_task_vector.py \
--alpha "vision=1.15" \
--out ./chai-ocr-vb-1.15
python merge_task_vector.py \
--alpha "vision=1.25" \
--out ./chai-ocr-vb-1.25
python merge_task_vector.py \
--alpha "vision=1.35" \
--out ./chai-ocr-vb-1.35
python merge_task_vector.py \
--alpha "vision=1.50" \
--out ./chai-ocr-vb-1.50
Results should be plotted against α for both the target task and regression-sensitive tasks.
Conceptually:
old_scans ↑
│
│ ●
│ ●
│ ●
│
└────────────────→ α
1.0 1.25 1.5
The optimal coefficient is not necessarily the largest one.
Task-vector extrapolation generally becomes increasingly unstable as parameters are moved farther outside the trained checkpoint.
Values around or above α ≈ 1.5 should therefore be treated as exploratory rather than assumed improvements.
If performance continues increasing at that point, the more defensible conclusion may be that the vision stack requires additional targeted training rather than stronger weight extrapolation.
Improving Degraded Scan Recognition
Task arithmetic is only one possible approach to improving old_scans.
Image preprocessing may provide a more predictable improvement for degraded documents.
Potential preprocessing stages include:
Input Scan
│
▼
Orientation Detection
│
▼
Deskew
│
▼
Denoising
│
▼
Local Contrast Enhancement
│
▼
Adaptive Binarization
│
▼
Resolution Normalization
│
▼
CHAI OCR
Useful preprocessing techniques include:
- deskewing;
- denoising;
- CLAHE contrast enhancement;
- adaptive thresholding;
- Sauvola binarization;
- background normalization;
- border removal;
- orientation correction; and
- intelligent resizing.
For heavily degraded scans, preprocessing should be benchmarked independently as well as in combination with CHAI OCR.
Intended Use
CHAI OCR is intended for research and experimentation involving:
- Optical Character Recognition;
- scanned-document transcription;
- document understanding;
- historical document processing;
- degraded document recognition;
- document layout interpretation;
- tables and forms;
- multi-column documents;
- visually structured text;
- document digitization; and
- multimodal document AI research.
Experimental Status
CHAI OCR should currently be considered an experimental research checkpoint.
The task-arithmetic modification has a plausible technical motivation, but the effect of α=1.25 must be established through controlled evaluation.
Users should therefore avoid assuming that:
larger α = better OCR
or that improvements on one benchmark category necessarily generalize to all document types.
Reproducibility
For reproducible comparisons, record at minimum:
Model revision
Benchmark revision
Transformers version
PyTorch version
CUDA version
GPU model
Image resolution
Image preprocessing
Prompt template
Generation parameters
Random seed
Batch size
Attention backend
Benchmark comparisons should not mix materially different preprocessing or decoding configurations.
Project Structure
A typical CHAI OCR repository may use the following structure:
chai-ocr/
│
├── README.md
├── config.json
├── generation_config.json
├── tokenizer.json
├── tokenizer_config.json
├── processor_config.json
├── preprocessor_config.json
├── chat_template.json
│
├── model-00001-of-000xx.safetensors
├── model-00002-of-000xx.safetensors
├── ...
├── model.safetensors.index.json
│
├── chai-ocr.png
│
├── merge_recipe.json
├── merge_task_vector.py
│
├── benchmark/
│ ├── olmocr/
│ ├── results/
│ └── evaluation_config.json
│
└── examples/
├── images/
├── documents/
└── inference.py
Summary
CHAI OCR explores a narrowly targeted modification to a Qwen 3.5 vision-language OCR checkpoint:
Qwen 3.5 Base
│
├───────────────┐
│ │
▼ ▼
Vision Tower Language Model
│ │
│ α = 1.0
│ │
α = 1.25 │
│ │
└───────┬───────┘
▼
CHAI OCR
The experiment is based on three observations:
- The vision tower changed substantially less than the language model during the source fine-tuning.
- Degraded scans were the source checkpoint's weakest measured OCR category.
- The vision tower is the component most directly responsible for extracting useful representations from degraded page images.
CHAI OCR therefore tests whether modest extrapolation of the learned vision task vector can improve degraded-document OCR while preserving the capabilities of the tuned language model.
The central criterion remains:
Improve difficult visual OCR without sacrificing the capabilities that already work well.
- Downloads last month
- 218
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="surendirakrishna/salad") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)