Visual Question Answering
PEFT
Safetensors
English
decision-model
calibration
typesafe
kev
jev
vision
qwen3.5
Instructions to use Jacqkues/sev-0.8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Jacqkues/sev-0.8b with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-0.8B-Base") model = PeftModel.from_pretrained(base_model, "Jacqkues/sev-0.8b") - Notebooks
- Google Colab
- Kaggle
File size: 10,390 Bytes
4a0a64e 958c183 c734ae0 4a0a64e 958c183 4a0a64e c734ae0 98d7a12 4a0a64e 958c183 4a0a64e 98d7a12 4a0a64e c734ae0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 | ---
language: en
license: other
license_name: research-use-derived-data
library_name: peft
base_model: jaredpalmer/kev-0.8b
base_model_relation: adapter
pipeline_tag: visual-question-answering
tags:
- decision-model
- calibration
- typesafe
- kev
- jev
- vision
- qwen3.5
datasets:
- Jacqkues/kev-vision-decisions-full
- Jacqkues/sev-ood
- Jacqkues/sev-generic-bench
---
# Sev-0.8B
**Sev** is a *decision model that reads images*: one image plus a state (text or JSON) and a set of typed questions go in,
a calibrated probability distribution per question comes out, in a single forward pass, with no text generation.
Question types are the TypeSafe `/v1/systemone` ones: `choice` (one of named options), `noul` (yes / no) and `score`
(ordered levels). New questions and options are given at request time; nothing is retrained.
It is [Kev-0.8B](https://huggingface.co/jaredpalmer/kev-0.8b) (a Jev-style text decision model: LoRA + pointer head on
`Qwen/Qwen3.5-0.8B-Base`) plugged into the **full Qwen3.5 vision-language model** and fine-tuned on
[kev-vision-decisions-full](https://huggingface.co/datasets/Jacqkues/kev-vision-decisions-full) (gated: research use, per-source licences). The vision tower is frozen;
the Kev LoRA (r = 16) and pointer head are trained further.
- **Status**: a research project, heavily vibe-coded with Claude Code. Not for production.
- **Document version**: [Sev-0.8B-docindex](https://huggingface.co/Jacqkues/sev-0.8b-docindex) (page indexing: 21 attributes)
- **Code**: `kev-vision` (`VisionDecisionModel`, training, dataset builder, evaluation)
- **Served temperature**: T = 1.32, fitted on the validation split (hard-label questions only)
## Use
```python
from kev.api import SystemOneRequest, to_record
from kev_vision.model import VisionDecisionModel
from PIL import Image
model = VisionDecisionModel("Jacqkues/sev-0.8b", device="cuda").eval() # Ampere/Ada GPU, MPS or CPU (see limits)
request = {"state": "A document page is attached.",
"questions": {"doc_type": {"type": "choice", "instructions": "What type of document is this?",
"criteria": {"invoice": "An invoice", "receipt": "A receipt", "letter": "A letter"}},
"handwritten": {"type": "noul", "instructions": "Is it mostly handwritten?"}}}
rec, meta = to_record(SystemOneRequest.model_validate(request)); rec["image"] = Image.open("page.jpg")
for m, p in zip(meta, model.probs(rec)):
print(m["id"], dict(zip(m["keys"], p.tolist())))
```
## How it works
```
[state] <vision_start> image tokens <vision_end> state text | [q] instructions [opt] option ... [decide]
Qwen3.5 (vision frozen, LoRA on the language model) β pointer head: <decide> Β· each option β softmax / T
```
Every question is its own causal row over the shared state (Qwen3.5 is a Gated DeltaNet hybrid), positions are Qwen's 3D
M-RoPE, and each option is scored at its end token. On text-only input Sev's code reproduces Kev exactly (max |Ξp| 5.7e-7).
## Training
- **Start**: `jaredpalmer/kev-0.8b` (adapter + head). A controlled comparison on 20k records picked it over a bare
Qwen3.5-0.8B-Base with a fresh LoRA of the same rank: better on 14 of 21 test sets, clearly on text (0.827 vs 0.766).
- **Data**: kev-vision-decisions-full, ~126k training records: VQAv2 yes/no + counting, A-OKVQA, ScienceQA, EuroSAT, Oxford
Pets, FairFace age, AVA aesthetics (soft targets), ChartNet, document types (DocLayNet, CORD receipts, invoices,
RVL-CDIP), control (Pong, Breakout, Freeway, Space Invaders, Ms. Pac-Man from JAT; LeRobot PushT, upweighted Γ2),
text replay from Kev's sources, and blank / noise images with a uniform target. Yes/no balanced per source, option
order randomised, images shared between train and held-out splits removed.
- **Recipe**: 1 epoch (136,502 passages), batch 8, AdamW lr 5e-5 OneCycle, bf16 autocast, Kev's augmentation (option
permutation, none-of-the-above, distractors) redrawn each pass. ~3 h on one A100 80 GB (Modal).
- **Validation**: loss 0.6635, accuracy 0.763, Brier 0.31.
## Results: test splits (200 records each, served temperature)
| test set | before (Kev-0.8B, zero-shot) | **Sev-0.8B** | Brier | ECE |
|---|---|---|---|---|
| `ai2d_ood` | 0.585 | **0.665** | 0.446 | 0.077 |
| `aokvqa` | 0.77 | **0.810** | 0.267 | 0.028 |
| `ava` | 0.23 | **0.640** | 0.706 | 0.397 |
| `breakout` | 0.37 | **0.445** | 0.650 | 0.068 |
| `chartnet` | 0.825 | **0.971** | 0.043 | 0.016 |
| `countbench_ood` | 0.555 | **0.730** | 0.399 | 0.106 |
| `documents` | 0.866 | **0.985** | 0.025 | 0.012 |
| `eurosat` | 0.588 | **0.940** | 0.087 | 0.029 |
| `fairface_age` | 0.295 | **0.595** | 0.544 | 0.059 |
| `freeway` | 0.862 | **0.814** | 0.359 | 0.176 |
| `mspacman` | 0.05 | **0.380** | 0.768 | 0.062 |
| `pets` | 0.815 | **0.975** | 0.036 | 0.025 |
| `pong` | 0.462 | **0.519** | 0.548 | 0.048 |
| `pope_ood` | 0.855 | **0.880** | 0.178 | 0.051 |
| `pusht` | 0.213 | **0.647** | 0.485 | 0.084 |
| `rvl_cdip` | 0.567 | **0.861** | 0.202 | 0.031 |
| `scienceqa` | 0.619 | **0.942** | 0.102 | 0.051 |
| `spaceinvaders` | 0.35 | **0.490** | 0.667 | 0.086 |
| `text` | 0.805 | **0.830** | 0.243 | 0.057 |
| `text_full` | 0.844 | **0.851** | 0.214 | 0.053 |
| `vqav2` | 0.796 | **0.853** | 0.210 | 0.042 |
Blank images (the image carries no answer): mean confidence **0.245** against a uniform 0.238.
## Results: sev-ood (never trained on)
Unseen Atari games, unseen classes, questions the image cannot answer, and corrupted images
([Jacqkues/sev-ood](https://huggingface.co/datasets/Jacqkues/sev-ood), gated, 500 records per set). *Confident errors* = wrong at
p β₯ 0.9; *coverage* = share of decisions that could be automated with at most 5% of them wrong.
| set | accuracy | ECE | confident errors | coverage at 5% error |
|---|---|---|---|---|
| `corrupted_eurosat` | **0.555** | 0.263 | 0.187 | 0.00 |
| `corrupted_pets` | **0.938** | 0.018 | 0.005 | 0.98 |
| `corrupted_vqav2` | **0.767** | 0.039 | 0.025 | 0.21 |
| `ood_beamrider` | **0.350** | 0.081 | 0.000 | 0.00 |
| `ood_boxing` | **0.162** | 0.038 | 0.000 | 0.00 |
| `ood_cars` | **0.872** | 0.054 | 0.000 | 0.75 |
| `ood_flowers` | **0.734** | 0.083 | 0.022 | 0.53 |
| `ood_food101` | **0.886** | 0.010 | 0.012 | 0.81 |
| `ood_qbert` | **0.278** | 0.011 | 0.000 | 0.00 |
| `ood_seaquest` | **0.160** | 0.027 | 0.000 | 0.00 |
| `unknowable` (uniform target) | mean confidence 0.396 vs chance 0.227 | | answered at β₯ 0.9: **1.2%** | |
## Results: generic vision benchmarks (never trained on)
Five public benchmarks, cast as typed questions: MMStar (A-D, vision-indispensable), MMBench-en dev (1,000 sampled),
MME (yes/no), RealWorldQA (A-F or yes/no) and HallusionBench (image split, yes/no). ScienceQA-sourced items were dropped
(Sev trained on ScienceQA). **Contamination control**: every benchmark image was hashed (256-bit average hash) against
every training image; the 305 of 6,001 items within 6 bits of a training image (mostly COCO photos via VQAv2 / A-OKVQA,
and ScienceQA) are removed below. Removing them changes no conclusion: they were slightly easier for all three models,
including the baseline that never saw them.
The baseline is **Qwen3.5-0.8B** (the post-trained VLM of the same family) in zero-shot, scored by its next-token
probability over the option letters (or Yes / No), so the same calibration metrics apply. All models see images capped at
448 x 448 pixels. Accuracy on the 5,696 clean items; gap = Sev minus Qwen, paired bootstrap 95% interval.
| benchmark (chance) | n | Qwen3.5-0.8B | **Sev-0.8B** | gap [95% CI] | ECE Qwen β Sev |
|---|---|---|---|---|---|
| MMStar (0.25) | 1,036 | 0.460 | **0.526** | +6.6 [+3.3, +9.9] | 0.18 β **0.08** |
| MMBench dev (0.26) | 927 | 0.774 | **0.825** | +5.1 [+2.6, +7.6] | **0.02** β 0.06 |
| MME (0.50) | 2,206 | 0.711 | **0.771** | +6.0 [+4.2, +7.9] | 0.12 β **0.01** |
| RealWorldQA (0.39) | 586 | 0.570 | **0.631** | +6.1 [+2.2, +10.2] | 0.15 β **0.05** |
| HallusionBench (0.50) | 941 | 0.644 | 0.629 | β1.5 [β4.8, +1.9] | 0.15 β **0.10** |
| **all** | 5,696 | 0.650 | **0.698** | **+4.7 [+3.5, +6.0]** | 0.12 β **0.02** |
Confident errors (wrong at p β₯ 0.9), all clean items: 4.5% for Qwen, **1.5%** for Sev. HallusionBench (visual illusions,
edited charts) is the exception: no model of this size is reliably better than chance-level guessing there.
## Results: docbench (document pages, never trained on)
DocLayNet test + validation pages, three yes/no questions (table of contents? figure? table?), 676 items, labels from
independent rules. See [Sev-0.8B-docindex](https://huggingface.co/Jacqkues/sev-0.8b-docindex) for the document model.
| model | accuracy | AUROC | ECE |
|---|---|---|---|
| Kev-0.8B | 0.886 | 0.934 | 0.236 |
| Sev-0.8B | 0.831 | 0.924 | **0.044** |
| Sev-0.8B-docindex | **0.928** | **0.976** | **0.018** |
## Limits
- **Control does not transfer**: unseen games stay near chance, and on trained games Sev is still below simply repeating
the previous action (Pong 0.52 vs ~0.69). It is cautious there (no confident errors), not competent.
- **Degraded satellite images**: on blurred / noised EuroSAT tiles Sev is over-confident (18.7% confident errors).
- **Aesthetics (AVA) is poorly calibrated** (ECE 0.40): a subjective target.
- Zero-shot classification of unseen classes is good (cars 0.87, food 0.89, flowers 0.73) but 1-4 points below the
20k-record Kev start: more task-specific training erodes a little general knowledge.
- **Hardware**: on Hopper GPUs (H100) flash-linear-attention refuses the Gated DeltaNet backward with torch 2.8's Triton
(incorrect results, fla #640): train on Ampere/Ada (A100, L4). On Macs the DeltaNet layers run a slow PyTorch fallback.
- It answers the questions it is asked; it does not verify identity, and it must not be used to infer sensitive
attributes (the training data asks apparent age only; gender and ethnicity were deliberately never asked).
## Licence
The adapter and head are derived from Kev-0.8B (Apache-2.0) and Qwen3.5 (Apache-2.0), but were trained on data that
includes research-only / non-commercial sources (RVL-CDIP, AVA, ScienceQA, AG News, Yelp; see the dataset card).
**Use for research only.** A commercially usable version would have to be retrained on permissively licensed sources only.
|