Visual Question Answering
PEFT
Safetensors
English
decision-model
calibration
typesafe
kev
jev
vision
qwen3.5
Jacqkues commited on
Commit
4a0a64e
·
verified ·
1 Parent(s): c394a01

Model card

Browse files
Files changed (1) hide show
  1. README.md +143 -0
README.md ADDED
@@ -0,0 +1,143 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language: en
3
+ license: other
4
+ license_name: research-use-derived-data
5
+ library_name: peft
6
+ base_model: jaredpalmer/kev-0.8b
7
+ base_model_relation: adapter
8
+ pipeline_tag: visual-question-answering
9
+ tags:
10
+ - decision-model
11
+ - calibration
12
+ - typesafe
13
+ - kev
14
+ - jev
15
+ - vision
16
+ - qwen3.5
17
+ datasets:
18
+ - Jacqkues/kev-vision-decisions-full
19
+ - Jacqkues/sev-ood
20
+ ---
21
+
22
+ # Sev-0.8B
23
+
24
+ **Sev** is a *decision model that reads images*: one image plus a state (text or JSON) and a set of typed questions go in,
25
+ a calibrated probability distribution per question comes out, in a single forward pass, with no text generation.
26
+ Question types are the TypeSafe `/v1/systemone` ones: `choice` (one of named options), `noul` (yes / no) and `score`
27
+ (ordered levels). New questions and options are given at request time; nothing is retrained.
28
+
29
+ It is [Kev-0.8B](https://huggingface.co/jaredpalmer/kev-0.8b) (a Jev-style text decision model: LoRA + pointer head on
30
+ `Qwen/Qwen3.5-0.8B-Base`) plugged into the **full Qwen3.5 vision-language model** and fine-tuned on
31
+ [kev-vision-decisions-full](https://huggingface.co/datasets/Jacqkues/kev-vision-decisions-full). The vision tower is frozen;
32
+ the Kev LoRA (r = 16) and pointer head are trained further.
33
+
34
+ - **Demo**: https://jacq-dumora--sev-demo-ui.modal.run (ID-photo check, document type, photo content, or your own questions)
35
+ - **Code**: `kev-vision` (`VisionDecisionModel`, training, dataset builder, evaluation)
36
+ - **Served temperature**: T = 1.32, fitted on the validation split (hard-label questions only)
37
+
38
+ ## Use
39
+
40
+ ```python
41
+ from kev.api import SystemOneRequest, to_record
42
+ from kev_vision.model import VisionDecisionModel
43
+ from PIL import Image
44
+
45
+ model = VisionDecisionModel("Jacqkues/sev-0.8b", device="cuda").eval() # Ampere/Ada GPU, MPS or CPU (see limits)
46
+ request = {"state": "A document page is attached.",
47
+ "questions": {"doc_type": {"type": "choice", "instructions": "What type of document is this?",
48
+ "criteria": {"invoice": "An invoice", "receipt": "A receipt", "letter": "A letter"}},
49
+ "handwritten": {"type": "noul", "instructions": "Is it mostly handwritten?"}}}
50
+ rec, meta = to_record(SystemOneRequest.model_validate(request)); rec["image"] = Image.open("page.jpg")
51
+ for m, p in zip(meta, model.probs(rec)):
52
+ print(m["id"], dict(zip(m["keys"], p.tolist())))
53
+ ```
54
+
55
+ ## How it works
56
+
57
+ ```
58
+ [state] <vision_start> image tokens <vision_end> state text | [q] instructions [opt] option ... [decide]
59
+ Qwen3.5 (vision frozen, LoRA on the language model) → pointer head: <decide> · each option → softmax / T
60
+ ```
61
+
62
+ Every question is its own causal row over the shared state (Qwen3.5 is a Gated DeltaNet hybrid), positions are Qwen's 3D
63
+ M-RoPE, and each option is scored at its end token. On text-only input Sev's code reproduces Kev exactly (max |Δp| 5.7e-7).
64
+
65
+ ## Training
66
+
67
+ - **Start**: `jaredpalmer/kev-0.8b` (adapter + head). A controlled comparison on 20k records picked it over a bare
68
+ Qwen3.5-0.8B-Base with a fresh LoRA of the same rank: better on 14 of 21 test sets, clearly on text (0.827 vs 0.766).
69
+ - **Data**: kev-vision-decisions-full, ~126k training records: VQAv2 yes/no + counting, A-OKVQA, ScienceQA, EuroSAT, Oxford
70
+ Pets, FairFace age, AVA aesthetics (soft targets), ChartNet, document types (DocLayNet, CORD receipts, invoices,
71
+ RVL-CDIP), control (Pong, Breakout, Freeway, Space Invaders, Ms. Pac-Man from JAT; LeRobot PushT, upweighted ×2),
72
+ text replay from Kev's sources, and blank / noise images with a uniform target. Yes/no balanced per source, option
73
+ order randomised, images shared between train and held-out splits removed.
74
+ - **Recipe**: 1 epoch (136,502 passages), batch 8, AdamW lr 5e-5 OneCycle, bf16 autocast, Kev's augmentation (option
75
+ permutation, none-of-the-above, distractors) redrawn each pass. ~3 h on one A100 80 GB (Modal).
76
+ - **Validation**: loss 0.6635, accuracy 0.763, Brier 0.31.
77
+
78
+ ## Results: test splits (200 records each, served temperature)
79
+
80
+ | test set | before (Kev-0.8B, zero-shot) | **Sev-0.8B** | Brier | ECE |
81
+ |---|---|---|---|---|
82
+ | `ai2d_ood` | 0.585 | **0.665** | 0.446 | 0.077 |
83
+ | `aokvqa` | 0.77 | **0.810** | 0.267 | 0.028 |
84
+ | `ava` | 0.23 | **0.640** | 0.706 | 0.397 |
85
+ | `breakout` | 0.37 | **0.445** | 0.650 | 0.068 |
86
+ | `chartnet` | 0.825 | **0.971** | 0.043 | 0.016 |
87
+ | `countbench_ood` | 0.555 | **0.730** | 0.399 | 0.106 |
88
+ | `documents` | 0.866 | **0.985** | 0.025 | 0.012 |
89
+ | `eurosat` | 0.588 | **0.940** | 0.087 | 0.029 |
90
+ | `fairface_age` | 0.295 | **0.595** | 0.544 | 0.059 |
91
+ | `freeway` | 0.862 | **0.814** | 0.359 | 0.176 |
92
+ | `mspacman` | 0.05 | **0.380** | 0.768 | 0.062 |
93
+ | `pets` | 0.815 | **0.975** | 0.036 | 0.025 |
94
+ | `pong` | 0.462 | **0.519** | 0.548 | 0.048 |
95
+ | `pope_ood` | 0.855 | **0.880** | 0.178 | 0.051 |
96
+ | `pusht` | 0.213 | **0.647** | 0.485 | 0.084 |
97
+ | `rvl_cdip` | 0.567 | **0.861** | 0.202 | 0.031 |
98
+ | `scienceqa` | 0.619 | **0.942** | 0.102 | 0.051 |
99
+ | `spaceinvaders` | 0.35 | **0.490** | 0.667 | 0.086 |
100
+ | `text` | 0.805 | **0.830** | 0.243 | 0.057 |
101
+ | `text_full` | 0.844 | **0.851** | 0.214 | 0.053 |
102
+ | `vqav2` | 0.796 | **0.853** | 0.210 | 0.042 |
103
+
104
+ Blank images (the image carries no answer): mean confidence **0.245** against a uniform 0.238.
105
+
106
+ ## Results: sev-ood (never trained on)
107
+
108
+ Unseen Atari games, unseen classes, questions the image cannot answer, and corrupted images
109
+ ([Jacqkues/sev-ood](https://huggingface.co/datasets/Jacqkues/sev-ood), 500 records per set). *Confident errors* = wrong at
110
+ p ≥ 0.9; *coverage* = share of decisions that could be automated with at most 5% of them wrong.
111
+
112
+ | set | accuracy | ECE | confident errors | coverage at 5% error |
113
+ |---|---|---|---|---|
114
+ | `corrupted_eurosat` | **0.555** | 0.263 | 0.187 | 0.00 |
115
+ | `corrupted_pets` | **0.938** | 0.018 | 0.005 | 0.98 |
116
+ | `corrupted_vqav2` | **0.767** | 0.039 | 0.025 | 0.21 |
117
+ | `ood_beamrider` | **0.350** | 0.081 | 0.000 | 0.00 |
118
+ | `ood_boxing` | **0.162** | 0.038 | 0.000 | 0.00 |
119
+ | `ood_cars` | **0.872** | 0.054 | 0.000 | 0.75 |
120
+ | `ood_flowers` | **0.734** | 0.083 | 0.022 | 0.53 |
121
+ | `ood_food101` | **0.886** | 0.010 | 0.012 | 0.81 |
122
+ | `ood_qbert` | **0.278** | 0.011 | 0.000 | 0.00 |
123
+ | `ood_seaquest` | **0.160** | 0.027 | 0.000 | 0.00 |
124
+ | `unknowable` (uniform target) | mean confidence 0.396 vs chance 0.227 | | answered at ≥ 0.9: **1.2%** | |
125
+
126
+ ## Limits
127
+
128
+ - **Control does not transfer**: unseen games stay near chance, and on trained games Sev is still below simply repeating
129
+ the previous action (Pong 0.52 vs ~0.69). It is cautious there (no confident errors), not competent.
130
+ - **Degraded satellite images**: on blurred / noised EuroSAT tiles Sev is over-confident (18.7% confident errors).
131
+ - **Aesthetics (AVA) is poorly calibrated** (ECE 0.40): a subjective target.
132
+ - Zero-shot classification of unseen classes is good (cars 0.87, food 0.89, flowers 0.73) but 1-4 points below the
133
+ 20k-record Kev start: more task-specific training erodes a little general knowledge.
134
+ - **Hardware**: on Hopper GPUs (H100) flash-linear-attention refuses the Gated DeltaNet backward with torch 2.8's Triton
135
+ (incorrect results, fla #640): train on Ampere/Ada (A100, L4). On Macs the DeltaNet layers run a slow PyTorch fallback.
136
+ - It answers the questions it is asked; it does not verify identity, and it must not be used to infer sensitive
137
+ attributes (the training data asks apparent age only; gender and ethnicity were deliberately never asked).
138
+
139
+ ## Licence
140
+
141
+ The adapter and head are derived from Kev-0.8B (Apache-2.0) and Qwen3.5 (Apache-2.0), but were trained on data that
142
+ includes research-only / non-commercial sources (RVL-CDIP, AVA, ScienceQA, AG News, Yelp; see the dataset card).
143
+ **Use for research only** unless you retrain on the public tier (`Jacqkues/kev-vision-decisions`).