Visual Question Answering
PEFT
Safetensors
English
decision-model
calibration
typesafe
kev
jev
vision
qwen3.5
Instructions to use Jacqkues/sev-0.8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Jacqkues/sev-0.8b with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-0.8B-Base") model = PeftModel.from_pretrained(base_model, "Jacqkues/sev-0.8b") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from Jacqkues/sev-0.8b: direct link, hf CLI and curl.
- Browser
- Download file 10.4 kB
-
https://huggingface.co/Jacqkues/sev-0.8b/resolve/main/README.md
- Command line
-
hf download hf://Jacqkues/sev-0.8b/README.md
-
curl -L -o README.md https://huggingface.co/Jacqkues/sev-0.8b/resolve/main/README.md
10.4 kB
| language: en | |
| license: other | |
| license_name: research-use-derived-data | |
| library_name: peft | |
| base_model: jaredpalmer/kev-0.8b | |
| base_model_relation: adapter | |
| pipeline_tag: visual-question-answering | |
| tags: | |
| - decision-model | |
| - calibration | |
| - typesafe | |
| - kev | |
| - jev | |
| - vision | |
| - qwen3.5 | |
| datasets: | |
| - Jacqkues/kev-vision-decisions-full | |
| - Jacqkues/sev-ood | |
| - Jacqkues/sev-generic-bench | |
| # Sev-0.8B | |
| **Sev** is a *decision model that reads images*: one image plus a state (text or JSON) and a set of typed questions go in, | |
| a calibrated probability distribution per question comes out, in a single forward pass, with no text generation. | |
| Question types are the TypeSafe `/v1/systemone` ones: `choice` (one of named options), `noul` (yes / no) and `score` | |
| (ordered levels). New questions and options are given at request time; nothing is retrained. | |
| It is [Kev-0.8B](https://huggingface.co/jaredpalmer/kev-0.8b) (a Jev-style text decision model: LoRA + pointer head on | |
| `Qwen/Qwen3.5-0.8B-Base`) plugged into the **full Qwen3.5 vision-language model** and fine-tuned on | |
| [kev-vision-decisions-full](https://huggingface.co/datasets/Jacqkues/kev-vision-decisions-full) (gated: research use, per-source licences). The vision tower is frozen; | |
| the Kev LoRA (r = 16) and pointer head are trained further. | |
| - **Status**: a research project, heavily vibe-coded with Claude Code. Not for production. | |
| - **Document version**: [Sev-0.8B-docindex](https://huggingface.co/Jacqkues/sev-0.8b-docindex) (page indexing: 21 attributes) | |
| - **Code**: `kev-vision` (`VisionDecisionModel`, training, dataset builder, evaluation) | |
| - **Served temperature**: T = 1.32, fitted on the validation split (hard-label questions only) | |
| ## Use | |
| ```python | |
| from kev.api import SystemOneRequest, to_record | |
| from kev_vision.model import VisionDecisionModel | |
| from PIL import Image | |
| model = VisionDecisionModel("Jacqkues/sev-0.8b", device="cuda").eval() # Ampere/Ada GPU, MPS or CPU (see limits) | |
| request = {"state": "A document page is attached.", | |
| "questions": {"doc_type": {"type": "choice", "instructions": "What type of document is this?", | |
| "criteria": {"invoice": "An invoice", "receipt": "A receipt", "letter": "A letter"}}, | |
| "handwritten": {"type": "noul", "instructions": "Is it mostly handwritten?"}}} | |
| rec, meta = to_record(SystemOneRequest.model_validate(request)); rec["image"] = Image.open("page.jpg") | |
| for m, p in zip(meta, model.probs(rec)): | |
| print(m["id"], dict(zip(m["keys"], p.tolist()))) | |
| ``` | |
| ## How it works | |
| ``` | |
| [state] <vision_start> image tokens <vision_end> state text | [q] instructions [opt] option ... [decide] | |
| Qwen3.5 (vision frozen, LoRA on the language model) β pointer head: <decide> Β· each option β softmax / T | |
| ``` | |
| Every question is its own causal row over the shared state (Qwen3.5 is a Gated DeltaNet hybrid), positions are Qwen's 3D | |
| M-RoPE, and each option is scored at its end token. On text-only input Sev's code reproduces Kev exactly (max |Ξp| 5.7e-7). | |
| ## Training | |
| - **Start**: `jaredpalmer/kev-0.8b` (adapter + head). A controlled comparison on 20k records picked it over a bare | |
| Qwen3.5-0.8B-Base with a fresh LoRA of the same rank: better on 14 of 21 test sets, clearly on text (0.827 vs 0.766). | |
| - **Data**: kev-vision-decisions-full, ~126k training records: VQAv2 yes/no + counting, A-OKVQA, ScienceQA, EuroSAT, Oxford | |
| Pets, FairFace age, AVA aesthetics (soft targets), ChartNet, document types (DocLayNet, CORD receipts, invoices, | |
| RVL-CDIP), control (Pong, Breakout, Freeway, Space Invaders, Ms. Pac-Man from JAT; LeRobot PushT, upweighted Γ2), | |
| text replay from Kev's sources, and blank / noise images with a uniform target. Yes/no balanced per source, option | |
| order randomised, images shared between train and held-out splits removed. | |
| - **Recipe**: 1 epoch (136,502 passages), batch 8, AdamW lr 5e-5 OneCycle, bf16 autocast, Kev's augmentation (option | |
| permutation, none-of-the-above, distractors) redrawn each pass. ~3 h on one A100 80 GB (Modal). | |
| - **Validation**: loss 0.6635, accuracy 0.763, Brier 0.31. | |
| ## Results: test splits (200 records each, served temperature) | |
| | test set | before (Kev-0.8B, zero-shot) | **Sev-0.8B** | Brier | ECE | | |
| |---|---|---|---|---| | |
| | `ai2d_ood` | 0.585 | **0.665** | 0.446 | 0.077 | | |
| | `aokvqa` | 0.77 | **0.810** | 0.267 | 0.028 | | |
| | `ava` | 0.23 | **0.640** | 0.706 | 0.397 | | |
| | `breakout` | 0.37 | **0.445** | 0.650 | 0.068 | | |
| | `chartnet` | 0.825 | **0.971** | 0.043 | 0.016 | | |
| | `countbench_ood` | 0.555 | **0.730** | 0.399 | 0.106 | | |
| | `documents` | 0.866 | **0.985** | 0.025 | 0.012 | | |
| | `eurosat` | 0.588 | **0.940** | 0.087 | 0.029 | | |
| | `fairface_age` | 0.295 | **0.595** | 0.544 | 0.059 | | |
| | `freeway` | 0.862 | **0.814** | 0.359 | 0.176 | | |
| | `mspacman` | 0.05 | **0.380** | 0.768 | 0.062 | | |
| | `pets` | 0.815 | **0.975** | 0.036 | 0.025 | | |
| | `pong` | 0.462 | **0.519** | 0.548 | 0.048 | | |
| | `pope_ood` | 0.855 | **0.880** | 0.178 | 0.051 | | |
| | `pusht` | 0.213 | **0.647** | 0.485 | 0.084 | | |
| | `rvl_cdip` | 0.567 | **0.861** | 0.202 | 0.031 | | |
| | `scienceqa` | 0.619 | **0.942** | 0.102 | 0.051 | | |
| | `spaceinvaders` | 0.35 | **0.490** | 0.667 | 0.086 | | |
| | `text` | 0.805 | **0.830** | 0.243 | 0.057 | | |
| | `text_full` | 0.844 | **0.851** | 0.214 | 0.053 | | |
| | `vqav2` | 0.796 | **0.853** | 0.210 | 0.042 | | |
| Blank images (the image carries no answer): mean confidence **0.245** against a uniform 0.238. | |
| ## Results: sev-ood (never trained on) | |
| Unseen Atari games, unseen classes, questions the image cannot answer, and corrupted images | |
| ([Jacqkues/sev-ood](https://huggingface.co/datasets/Jacqkues/sev-ood), gated, 500 records per set). *Confident errors* = wrong at | |
| p β₯ 0.9; *coverage* = share of decisions that could be automated with at most 5% of them wrong. | |
| | set | accuracy | ECE | confident errors | coverage at 5% error | | |
| |---|---|---|---|---| | |
| | `corrupted_eurosat` | **0.555** | 0.263 | 0.187 | 0.00 | | |
| | `corrupted_pets` | **0.938** | 0.018 | 0.005 | 0.98 | | |
| | `corrupted_vqav2` | **0.767** | 0.039 | 0.025 | 0.21 | | |
| | `ood_beamrider` | **0.350** | 0.081 | 0.000 | 0.00 | | |
| | `ood_boxing` | **0.162** | 0.038 | 0.000 | 0.00 | | |
| | `ood_cars` | **0.872** | 0.054 | 0.000 | 0.75 | | |
| | `ood_flowers` | **0.734** | 0.083 | 0.022 | 0.53 | | |
| | `ood_food101` | **0.886** | 0.010 | 0.012 | 0.81 | | |
| | `ood_qbert` | **0.278** | 0.011 | 0.000 | 0.00 | | |
| | `ood_seaquest` | **0.160** | 0.027 | 0.000 | 0.00 | | |
| | `unknowable` (uniform target) | mean confidence 0.396 vs chance 0.227 | | answered at β₯ 0.9: **1.2%** | | | |
| ## Results: generic vision benchmarks (never trained on) | |
| Five public benchmarks, cast as typed questions: MMStar (A-D, vision-indispensable), MMBench-en dev (1,000 sampled), | |
| MME (yes/no), RealWorldQA (A-F or yes/no) and HallusionBench (image split, yes/no). ScienceQA-sourced items were dropped | |
| (Sev trained on ScienceQA). **Contamination control**: every benchmark image was hashed (256-bit average hash) against | |
| every training image; the 305 of 6,001 items within 6 bits of a training image (mostly COCO photos via VQAv2 / A-OKVQA, | |
| and ScienceQA) are removed below. Removing them changes no conclusion: they were slightly easier for all three models, | |
| including the baseline that never saw them. | |
| The baseline is **Qwen3.5-0.8B** (the post-trained VLM of the same family) in zero-shot, scored by its next-token | |
| probability over the option letters (or Yes / No), so the same calibration metrics apply. All models see images capped at | |
| 448 x 448 pixels. Accuracy on the 5,696 clean items; gap = Sev minus Qwen, paired bootstrap 95% interval. | |
| | benchmark (chance) | n | Qwen3.5-0.8B | **Sev-0.8B** | gap [95% CI] | ECE Qwen β Sev | | |
| |---|---|---|---|---|---| | |
| | MMStar (0.25) | 1,036 | 0.460 | **0.526** | +6.6 [+3.3, +9.9] | 0.18 β **0.08** | | |
| | MMBench dev (0.26) | 927 | 0.774 | **0.825** | +5.1 [+2.6, +7.6] | **0.02** β 0.06 | | |
| | MME (0.50) | 2,206 | 0.711 | **0.771** | +6.0 [+4.2, +7.9] | 0.12 β **0.01** | | |
| | RealWorldQA (0.39) | 586 | 0.570 | **0.631** | +6.1 [+2.2, +10.2] | 0.15 β **0.05** | | |
| | HallusionBench (0.50) | 941 | 0.644 | 0.629 | β1.5 [β4.8, +1.9] | 0.15 β **0.10** | | |
| | **all** | 5,696 | 0.650 | **0.698** | **+4.7 [+3.5, +6.0]** | 0.12 β **0.02** | | |
| Confident errors (wrong at p β₯ 0.9), all clean items: 4.5% for Qwen, **1.5%** for Sev. HallusionBench (visual illusions, | |
| edited charts) is the exception: no model of this size is reliably better than chance-level guessing there. | |
| ## Results: docbench (document pages, never trained on) | |
| DocLayNet test + validation pages, three yes/no questions (table of contents? figure? table?), 676 items, labels from | |
| independent rules. See [Sev-0.8B-docindex](https://huggingface.co/Jacqkues/sev-0.8b-docindex) for the document model. | |
| | model | accuracy | AUROC | ECE | | |
| |---|---|---|---| | |
| | Kev-0.8B | 0.886 | 0.934 | 0.236 | | |
| | Sev-0.8B | 0.831 | 0.924 | **0.044** | | |
| | Sev-0.8B-docindex | **0.928** | **0.976** | **0.018** | | |
| ## Limits | |
| - **Control does not transfer**: unseen games stay near chance, and on trained games Sev is still below simply repeating | |
| the previous action (Pong 0.52 vs ~0.69). It is cautious there (no confident errors), not competent. | |
| - **Degraded satellite images**: on blurred / noised EuroSAT tiles Sev is over-confident (18.7% confident errors). | |
| - **Aesthetics (AVA) is poorly calibrated** (ECE 0.40): a subjective target. | |
| - Zero-shot classification of unseen classes is good (cars 0.87, food 0.89, flowers 0.73) but 1-4 points below the | |
| 20k-record Kev start: more task-specific training erodes a little general knowledge. | |
| - **Hardware**: on Hopper GPUs (H100) flash-linear-attention refuses the Gated DeltaNet backward with torch 2.8's Triton | |
| (incorrect results, fla #640): train on Ampere/Ada (A100, L4). On Macs the DeltaNet layers run a slow PyTorch fallback. | |
| - It answers the questions it is asked; it does not verify identity, and it must not be used to infer sensitive | |
| attributes (the training data asks apparent age only; gender and ethnicity were deliberately never asked). | |
| ## Licence | |
| The adapter and head are derived from Kev-0.8B (Apache-2.0) and Qwen3.5 (Apache-2.0), but were trained on data that | |
| includes research-only / non-commercial sources (RVL-CDIP, AVA, ScienceQA, AG News, Yelp; see the dataset card). | |
| **Use for research only.** A commercially usable version would have to be retrained on permissively licensed sources only. | |