--- license: apache-2.0 base_model: Qwen/Qwen3-VL-4B-Instruct pipeline_tag: image-text-to-text tags: - document-understanding - information-extraction - vision-language - qwen3-vl library_name: transformers --- # obj_v1 Vision-language model fine-tuned for structured data extraction from Indian financial documents. Give it a page image and a JSON schema; it returns the schema filled in from what is on the page. A 4B vision-language model, LoRA fine-tuned and merged ## Authors

   Rachit Kumar on LinkedIn    Ahmed Zaweel on LinkedIn

## Serving with vLLM ```bash vllm serve objectai/obj_v1 \ --served-model-name obj_v1 \ --max-model-len 16384 \ --limit-mm-per-prompt '{"image":1}' \ --mm-processor-kwargs '{"max_pixels":1003520}' \ --trust-remote-code ``` `max_pixels` is 1280x28x28, the resolution the model was trained at. Raising it wastes KV cache; lowering it makes small print unreadable. ## Calling it The server is OpenAI-compatible, so an ordinary chat completion works: ```python import base64, json, openai client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") image = base64.b64encode(open("cheque.jpg", "rb").read()).decode() schema = {"cheque_details": {"amount": "number", "payee": "string", "date": "string", "cheque_number": "string"}} response = client.chat.completions.create( model="obj_v1", temperature=0.0, max_tokens=8192, messages=[ {"role": "system", "content": "You are a document data extraction model. " "Extract only values present in the document. " "Use null for fields that are absent or illegible. " "Output a single compact JSON object matching the requested schema. " "No prose, no markdown, no explanation."}, {"role": "user", "content": [ {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image}"}}, {"type": "text", "text": f"document_type: cheque\nschema: {json.dumps(schema)}"}, ]}, ], ) print(response.choices[0].message.content) ``` ## Prompt format Match training or accuracy drops. The system prompt above is verbatim, and the user turn is the image followed by exactly two lines: ``` document_type: schema: ``` Set `temperature=0.0` so the same page yields the same answer. ## Requirements | | | | --- | --- | | Weights | 8.9 GB (bf16) | | VRAM | 16 GB minimum, 24 GB comfortable | | Precision | bf16 (Ampere or newer; use fp16 below that) | | Context | 16384 covers the longest documents | Runs on an L4, A10G, L40S, A100 or RTX 4090. On a T4 add `--dtype float16`. Long documents matter: `bank_statement` and `form16` answers run to ~2500 tokens, so `max_tokens` below 4096 truncates them mid-JSON. ## Output Compact JSON matching the requested schema. Fields absent from the page come back `null` rather than guessed. Values found on the page that the schema did not ask for are placed under `extras` when that key is included in the schema. ## Limitations - Trained on Indian financial documents; other domains and layouts are untested. - Handwriting is the weakest case, particularly digits at low resolution. - The model does not verify its own arithmetic. Totals that must reconcile should be checked by the caller. ## License Apache 2.0. Fine-tuned from Qwen3-VL-4B-Instruct, which is Apache 2.0.