Visual Question Answering
PEFT
Safetensors
English
decision-model
calibration
typesafe
kev
jev
vision
qwen3.5
File size: 10,390 Bytes
4a0a64e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
958c183
 
c734ae0
4a0a64e
 
 
 
 
 
 
 
 
 
 
958c183
4a0a64e
 
c734ae0
98d7a12
4a0a64e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
958c183
4a0a64e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
98d7a12
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4a0a64e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c734ae0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
---
language: en
license: other
license_name: research-use-derived-data
library_name: peft
base_model: jaredpalmer/kev-0.8b
base_model_relation: adapter
pipeline_tag: visual-question-answering
tags:
  - decision-model
  - calibration
  - typesafe
  - kev
  - jev
  - vision
  - qwen3.5
datasets:
  - Jacqkues/kev-vision-decisions-full
  - Jacqkues/sev-ood
  - Jacqkues/sev-generic-bench
---

# Sev-0.8B

**Sev** is a *decision model that reads images*: one image plus a state (text or JSON) and a set of typed questions go in,
a calibrated probability distribution per question comes out, in a single forward pass, with no text generation.
Question types are the TypeSafe `/v1/systemone` ones: `choice` (one of named options), `noul` (yes / no) and `score`
(ordered levels). New questions and options are given at request time; nothing is retrained.

It is [Kev-0.8B](https://huggingface.co/jaredpalmer/kev-0.8b) (a Jev-style text decision model: LoRA + pointer head on
`Qwen/Qwen3.5-0.8B-Base`) plugged into the **full Qwen3.5 vision-language model** and fine-tuned on
[kev-vision-decisions-full](https://huggingface.co/datasets/Jacqkues/kev-vision-decisions-full) (gated: research use, per-source licences). The vision tower is frozen;
the Kev LoRA (r = 16) and pointer head are trained further.

- **Status**: a research project, heavily vibe-coded with Claude Code. Not for production.
- **Document version**: [Sev-0.8B-docindex](https://huggingface.co/Jacqkues/sev-0.8b-docindex) (page indexing: 21 attributes)
- **Code**: `kev-vision` (`VisionDecisionModel`, training, dataset builder, evaluation)
- **Served temperature**: T = 1.32, fitted on the validation split (hard-label questions only)

## Use

```python
from kev.api import SystemOneRequest, to_record
from kev_vision.model import VisionDecisionModel
from PIL import Image

model = VisionDecisionModel("Jacqkues/sev-0.8b", device="cuda").eval()     # Ampere/Ada GPU, MPS or CPU (see limits)
request = {"state": "A document page is attached.",
           "questions": {"doc_type": {"type": "choice", "instructions": "What type of document is this?",
                                      "criteria": {"invoice": "An invoice", "receipt": "A receipt", "letter": "A letter"}},
                         "handwritten": {"type": "noul", "instructions": "Is it mostly handwritten?"}}}
rec, meta = to_record(SystemOneRequest.model_validate(request)); rec["image"] = Image.open("page.jpg")
for m, p in zip(meta, model.probs(rec)):
    print(m["id"], dict(zip(m["keys"], p.tolist())))
```

## How it works

```
[state] <vision_start> image tokens <vision_end> state text  |  [q] instructions [opt] option ... [decide]
Qwen3.5 (vision frozen, LoRA on the language model)          β†’  pointer head: <decide> Β· each option  β†’  softmax / T
```

Every question is its own causal row over the shared state (Qwen3.5 is a Gated DeltaNet hybrid), positions are Qwen's 3D
M-RoPE, and each option is scored at its end token. On text-only input Sev's code reproduces Kev exactly (max |Ξ”p| 5.7e-7).

## Training

- **Start**: `jaredpalmer/kev-0.8b` (adapter + head). A controlled comparison on 20k records picked it over a bare
  Qwen3.5-0.8B-Base with a fresh LoRA of the same rank: better on 14 of 21 test sets, clearly on text (0.827 vs 0.766).
- **Data**: kev-vision-decisions-full, ~126k training records: VQAv2 yes/no + counting, A-OKVQA, ScienceQA, EuroSAT, Oxford
  Pets, FairFace age, AVA aesthetics (soft targets), ChartNet, document types (DocLayNet, CORD receipts, invoices,
  RVL-CDIP), control (Pong, Breakout, Freeway, Space Invaders, Ms. Pac-Man from JAT; LeRobot PushT, upweighted Γ—2),
  text replay from Kev's sources, and blank / noise images with a uniform target. Yes/no balanced per source, option
  order randomised, images shared between train and held-out splits removed.
- **Recipe**: 1 epoch (136,502 passages), batch 8, AdamW lr 5e-5 OneCycle, bf16 autocast, Kev's augmentation (option
  permutation, none-of-the-above, distractors) redrawn each pass. ~3 h on one A100 80 GB (Modal).
- **Validation**: loss 0.6635, accuracy 0.763, Brier 0.31.

## Results: test splits (200 records each, served temperature)

| test set | before (Kev-0.8B, zero-shot) | **Sev-0.8B** | Brier | ECE |
|---|---|---|---|---|
| `ai2d_ood` | 0.585 | **0.665** | 0.446 | 0.077 |
| `aokvqa` | 0.77 | **0.810** | 0.267 | 0.028 |
| `ava` | 0.23 | **0.640** | 0.706 | 0.397 |
| `breakout` | 0.37 | **0.445** | 0.650 | 0.068 |
| `chartnet` | 0.825 | **0.971** | 0.043 | 0.016 |
| `countbench_ood` | 0.555 | **0.730** | 0.399 | 0.106 |
| `documents` | 0.866 | **0.985** | 0.025 | 0.012 |
| `eurosat` | 0.588 | **0.940** | 0.087 | 0.029 |
| `fairface_age` | 0.295 | **0.595** | 0.544 | 0.059 |
| `freeway` | 0.862 | **0.814** | 0.359 | 0.176 |
| `mspacman` | 0.05 | **0.380** | 0.768 | 0.062 |
| `pets` | 0.815 | **0.975** | 0.036 | 0.025 |
| `pong` | 0.462 | **0.519** | 0.548 | 0.048 |
| `pope_ood` | 0.855 | **0.880** | 0.178 | 0.051 |
| `pusht` | 0.213 | **0.647** | 0.485 | 0.084 |
| `rvl_cdip` | 0.567 | **0.861** | 0.202 | 0.031 |
| `scienceqa` | 0.619 | **0.942** | 0.102 | 0.051 |
| `spaceinvaders` | 0.35 | **0.490** | 0.667 | 0.086 |
| `text` | 0.805 | **0.830** | 0.243 | 0.057 |
| `text_full` | 0.844 | **0.851** | 0.214 | 0.053 |
| `vqav2` | 0.796 | **0.853** | 0.210 | 0.042 |

Blank images (the image carries no answer): mean confidence **0.245** against a uniform 0.238.

## Results: sev-ood (never trained on)

Unseen Atari games, unseen classes, questions the image cannot answer, and corrupted images
([Jacqkues/sev-ood](https://huggingface.co/datasets/Jacqkues/sev-ood), gated, 500 records per set). *Confident errors* = wrong at
p β‰₯ 0.9; *coverage* = share of decisions that could be automated with at most 5% of them wrong.

| set | accuracy | ECE | confident errors | coverage at 5% error |
|---|---|---|---|---|
| `corrupted_eurosat` | **0.555** | 0.263 | 0.187 | 0.00 |
| `corrupted_pets` | **0.938** | 0.018 | 0.005 | 0.98 |
| `corrupted_vqav2` | **0.767** | 0.039 | 0.025 | 0.21 |
| `ood_beamrider` | **0.350** | 0.081 | 0.000 | 0.00 |
| `ood_boxing` | **0.162** | 0.038 | 0.000 | 0.00 |
| `ood_cars` | **0.872** | 0.054 | 0.000 | 0.75 |
| `ood_flowers` | **0.734** | 0.083 | 0.022 | 0.53 |
| `ood_food101` | **0.886** | 0.010 | 0.012 | 0.81 |
| `ood_qbert` | **0.278** | 0.011 | 0.000 | 0.00 |
| `ood_seaquest` | **0.160** | 0.027 | 0.000 | 0.00 |
| `unknowable` (uniform target) | mean confidence 0.396 vs chance 0.227 | | answered at β‰₯ 0.9: **1.2%** | |

## Results: generic vision benchmarks (never trained on)

Five public benchmarks, cast as typed questions: MMStar (A-D, vision-indispensable), MMBench-en dev (1,000 sampled),
MME (yes/no), RealWorldQA (A-F or yes/no) and HallusionBench (image split, yes/no). ScienceQA-sourced items were dropped
(Sev trained on ScienceQA). **Contamination control**: every benchmark image was hashed (256-bit average hash) against
every training image; the 305 of 6,001 items within 6 bits of a training image (mostly COCO photos via VQAv2 / A-OKVQA,
and ScienceQA) are removed below. Removing them changes no conclusion: they were slightly easier for all three models,
including the baseline that never saw them.

The baseline is **Qwen3.5-0.8B** (the post-trained VLM of the same family) in zero-shot, scored by its next-token
probability over the option letters (or Yes / No), so the same calibration metrics apply. All models see images capped at
448 x 448 pixels. Accuracy on the 5,696 clean items; gap = Sev minus Qwen, paired bootstrap 95% interval.

| benchmark (chance) | n | Qwen3.5-0.8B | **Sev-0.8B** | gap [95% CI] | ECE Qwen β†’ Sev |
|---|---|---|---|---|---|
| MMStar (0.25) | 1,036 | 0.460 | **0.526** | +6.6 [+3.3, +9.9] | 0.18 β†’ **0.08** |
| MMBench dev (0.26) | 927 | 0.774 | **0.825** | +5.1 [+2.6, +7.6] | **0.02** β†’ 0.06 |
| MME (0.50) | 2,206 | 0.711 | **0.771** | +6.0 [+4.2, +7.9] | 0.12 β†’ **0.01** |
| RealWorldQA (0.39) | 586 | 0.570 | **0.631** | +6.1 [+2.2, +10.2] | 0.15 β†’ **0.05** |
| HallusionBench (0.50) | 941 | 0.644 | 0.629 | βˆ’1.5 [βˆ’4.8, +1.9] | 0.15 β†’ **0.10** |
| **all** | 5,696 | 0.650 | **0.698** | **+4.7 [+3.5, +6.0]** | 0.12 β†’ **0.02** |

Confident errors (wrong at p β‰₯ 0.9), all clean items: 4.5% for Qwen, **1.5%** for Sev. HallusionBench (visual illusions,
edited charts) is the exception: no model of this size is reliably better than chance-level guessing there.

## Results: docbench (document pages, never trained on)

DocLayNet test + validation pages, three yes/no questions (table of contents? figure? table?), 676 items, labels from
independent rules. See [Sev-0.8B-docindex](https://huggingface.co/Jacqkues/sev-0.8b-docindex) for the document model.

| model | accuracy | AUROC | ECE |
|---|---|---|---|
| Kev-0.8B | 0.886 | 0.934 | 0.236 |
| Sev-0.8B | 0.831 | 0.924 | **0.044** |
| Sev-0.8B-docindex | **0.928** | **0.976** | **0.018** |

## Limits

- **Control does not transfer**: unseen games stay near chance, and on trained games Sev is still below simply repeating
  the previous action (Pong 0.52 vs ~0.69). It is cautious there (no confident errors), not competent.
- **Degraded satellite images**: on blurred / noised EuroSAT tiles Sev is over-confident (18.7% confident errors).
- **Aesthetics (AVA) is poorly calibrated** (ECE 0.40): a subjective target.
- Zero-shot classification of unseen classes is good (cars 0.87, food 0.89, flowers 0.73) but 1-4 points below the
  20k-record Kev start: more task-specific training erodes a little general knowledge.
- **Hardware**: on Hopper GPUs (H100) flash-linear-attention refuses the Gated DeltaNet backward with torch 2.8's Triton
  (incorrect results, fla #640): train on Ampere/Ada (A100, L4). On Macs the DeltaNet layers run a slow PyTorch fallback.
- It answers the questions it is asked; it does not verify identity, and it must not be used to infer sensitive
  attributes (the training data asks apparent age only; gender and ethnicity were deliberately never asked).

## Licence

The adapter and head are derived from Kev-0.8B (Apache-2.0) and Qwen3.5 (Apache-2.0), but were trained on data that
includes research-only / non-commercial sources (RVL-CDIP, AVA, ScienceQA, AG News, Yelp; see the dataset card).
**Use for research only.** A commercially usable version would have to be retrained on permissively licensed sources only.