Visual Question Answering
PEFT
Safetensors
English
decision-model
calibration
typesafe
kev
jev
vision
qwen3.5
Instructions to use Jacqkues/sev-0.8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Jacqkues/sev-0.8b with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-0.8B-Base") model = PeftModel.from_pretrained(base_model, "Jacqkues/sev-0.8b") - Notebooks
- Google Colab
- Kaggle
Model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,143 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language: en
|
| 3 |
+
license: other
|
| 4 |
+
license_name: research-use-derived-data
|
| 5 |
+
library_name: peft
|
| 6 |
+
base_model: jaredpalmer/kev-0.8b
|
| 7 |
+
base_model_relation: adapter
|
| 8 |
+
pipeline_tag: visual-question-answering
|
| 9 |
+
tags:
|
| 10 |
+
- decision-model
|
| 11 |
+
- calibration
|
| 12 |
+
- typesafe
|
| 13 |
+
- kev
|
| 14 |
+
- jev
|
| 15 |
+
- vision
|
| 16 |
+
- qwen3.5
|
| 17 |
+
datasets:
|
| 18 |
+
- Jacqkues/kev-vision-decisions-full
|
| 19 |
+
- Jacqkues/sev-ood
|
| 20 |
+
---
|
| 21 |
+
|
| 22 |
+
# Sev-0.8B
|
| 23 |
+
|
| 24 |
+
**Sev** is a *decision model that reads images*: one image plus a state (text or JSON) and a set of typed questions go in,
|
| 25 |
+
a calibrated probability distribution per question comes out, in a single forward pass, with no text generation.
|
| 26 |
+
Question types are the TypeSafe `/v1/systemone` ones: `choice` (one of named options), `noul` (yes / no) and `score`
|
| 27 |
+
(ordered levels). New questions and options are given at request time; nothing is retrained.
|
| 28 |
+
|
| 29 |
+
It is [Kev-0.8B](https://huggingface.co/jaredpalmer/kev-0.8b) (a Jev-style text decision model: LoRA + pointer head on
|
| 30 |
+
`Qwen/Qwen3.5-0.8B-Base`) plugged into the **full Qwen3.5 vision-language model** and fine-tuned on
|
| 31 |
+
[kev-vision-decisions-full](https://huggingface.co/datasets/Jacqkues/kev-vision-decisions-full). The vision tower is frozen;
|
| 32 |
+
the Kev LoRA (r = 16) and pointer head are trained further.
|
| 33 |
+
|
| 34 |
+
- **Demo**: https://jacq-dumora--sev-demo-ui.modal.run (ID-photo check, document type, photo content, or your own questions)
|
| 35 |
+
- **Code**: `kev-vision` (`VisionDecisionModel`, training, dataset builder, evaluation)
|
| 36 |
+
- **Served temperature**: T = 1.32, fitted on the validation split (hard-label questions only)
|
| 37 |
+
|
| 38 |
+
## Use
|
| 39 |
+
|
| 40 |
+
```python
|
| 41 |
+
from kev.api import SystemOneRequest, to_record
|
| 42 |
+
from kev_vision.model import VisionDecisionModel
|
| 43 |
+
from PIL import Image
|
| 44 |
+
|
| 45 |
+
model = VisionDecisionModel("Jacqkues/sev-0.8b", device="cuda").eval() # Ampere/Ada GPU, MPS or CPU (see limits)
|
| 46 |
+
request = {"state": "A document page is attached.",
|
| 47 |
+
"questions": {"doc_type": {"type": "choice", "instructions": "What type of document is this?",
|
| 48 |
+
"criteria": {"invoice": "An invoice", "receipt": "A receipt", "letter": "A letter"}},
|
| 49 |
+
"handwritten": {"type": "noul", "instructions": "Is it mostly handwritten?"}}}
|
| 50 |
+
rec, meta = to_record(SystemOneRequest.model_validate(request)); rec["image"] = Image.open("page.jpg")
|
| 51 |
+
for m, p in zip(meta, model.probs(rec)):
|
| 52 |
+
print(m["id"], dict(zip(m["keys"], p.tolist())))
|
| 53 |
+
```
|
| 54 |
+
|
| 55 |
+
## How it works
|
| 56 |
+
|
| 57 |
+
```
|
| 58 |
+
[state] <vision_start> image tokens <vision_end> state text | [q] instructions [opt] option ... [decide]
|
| 59 |
+
Qwen3.5 (vision frozen, LoRA on the language model) → pointer head: <decide> · each option → softmax / T
|
| 60 |
+
```
|
| 61 |
+
|
| 62 |
+
Every question is its own causal row over the shared state (Qwen3.5 is a Gated DeltaNet hybrid), positions are Qwen's 3D
|
| 63 |
+
M-RoPE, and each option is scored at its end token. On text-only input Sev's code reproduces Kev exactly (max |Δp| 5.7e-7).
|
| 64 |
+
|
| 65 |
+
## Training
|
| 66 |
+
|
| 67 |
+
- **Start**: `jaredpalmer/kev-0.8b` (adapter + head). A controlled comparison on 20k records picked it over a bare
|
| 68 |
+
Qwen3.5-0.8B-Base with a fresh LoRA of the same rank: better on 14 of 21 test sets, clearly on text (0.827 vs 0.766).
|
| 69 |
+
- **Data**: kev-vision-decisions-full, ~126k training records: VQAv2 yes/no + counting, A-OKVQA, ScienceQA, EuroSAT, Oxford
|
| 70 |
+
Pets, FairFace age, AVA aesthetics (soft targets), ChartNet, document types (DocLayNet, CORD receipts, invoices,
|
| 71 |
+
RVL-CDIP), control (Pong, Breakout, Freeway, Space Invaders, Ms. Pac-Man from JAT; LeRobot PushT, upweighted ×2),
|
| 72 |
+
text replay from Kev's sources, and blank / noise images with a uniform target. Yes/no balanced per source, option
|
| 73 |
+
order randomised, images shared between train and held-out splits removed.
|
| 74 |
+
- **Recipe**: 1 epoch (136,502 passages), batch 8, AdamW lr 5e-5 OneCycle, bf16 autocast, Kev's augmentation (option
|
| 75 |
+
permutation, none-of-the-above, distractors) redrawn each pass. ~3 h on one A100 80 GB (Modal).
|
| 76 |
+
- **Validation**: loss 0.6635, accuracy 0.763, Brier 0.31.
|
| 77 |
+
|
| 78 |
+
## Results: test splits (200 records each, served temperature)
|
| 79 |
+
|
| 80 |
+
| test set | before (Kev-0.8B, zero-shot) | **Sev-0.8B** | Brier | ECE |
|
| 81 |
+
|---|---|---|---|---|
|
| 82 |
+
| `ai2d_ood` | 0.585 | **0.665** | 0.446 | 0.077 |
|
| 83 |
+
| `aokvqa` | 0.77 | **0.810** | 0.267 | 0.028 |
|
| 84 |
+
| `ava` | 0.23 | **0.640** | 0.706 | 0.397 |
|
| 85 |
+
| `breakout` | 0.37 | **0.445** | 0.650 | 0.068 |
|
| 86 |
+
| `chartnet` | 0.825 | **0.971** | 0.043 | 0.016 |
|
| 87 |
+
| `countbench_ood` | 0.555 | **0.730** | 0.399 | 0.106 |
|
| 88 |
+
| `documents` | 0.866 | **0.985** | 0.025 | 0.012 |
|
| 89 |
+
| `eurosat` | 0.588 | **0.940** | 0.087 | 0.029 |
|
| 90 |
+
| `fairface_age` | 0.295 | **0.595** | 0.544 | 0.059 |
|
| 91 |
+
| `freeway` | 0.862 | **0.814** | 0.359 | 0.176 |
|
| 92 |
+
| `mspacman` | 0.05 | **0.380** | 0.768 | 0.062 |
|
| 93 |
+
| `pets` | 0.815 | **0.975** | 0.036 | 0.025 |
|
| 94 |
+
| `pong` | 0.462 | **0.519** | 0.548 | 0.048 |
|
| 95 |
+
| `pope_ood` | 0.855 | **0.880** | 0.178 | 0.051 |
|
| 96 |
+
| `pusht` | 0.213 | **0.647** | 0.485 | 0.084 |
|
| 97 |
+
| `rvl_cdip` | 0.567 | **0.861** | 0.202 | 0.031 |
|
| 98 |
+
| `scienceqa` | 0.619 | **0.942** | 0.102 | 0.051 |
|
| 99 |
+
| `spaceinvaders` | 0.35 | **0.490** | 0.667 | 0.086 |
|
| 100 |
+
| `text` | 0.805 | **0.830** | 0.243 | 0.057 |
|
| 101 |
+
| `text_full` | 0.844 | **0.851** | 0.214 | 0.053 |
|
| 102 |
+
| `vqav2` | 0.796 | **0.853** | 0.210 | 0.042 |
|
| 103 |
+
|
| 104 |
+
Blank images (the image carries no answer): mean confidence **0.245** against a uniform 0.238.
|
| 105 |
+
|
| 106 |
+
## Results: sev-ood (never trained on)
|
| 107 |
+
|
| 108 |
+
Unseen Atari games, unseen classes, questions the image cannot answer, and corrupted images
|
| 109 |
+
([Jacqkues/sev-ood](https://huggingface.co/datasets/Jacqkues/sev-ood), 500 records per set). *Confident errors* = wrong at
|
| 110 |
+
p ≥ 0.9; *coverage* = share of decisions that could be automated with at most 5% of them wrong.
|
| 111 |
+
|
| 112 |
+
| set | accuracy | ECE | confident errors | coverage at 5% error |
|
| 113 |
+
|---|---|---|---|---|
|
| 114 |
+
| `corrupted_eurosat` | **0.555** | 0.263 | 0.187 | 0.00 |
|
| 115 |
+
| `corrupted_pets` | **0.938** | 0.018 | 0.005 | 0.98 |
|
| 116 |
+
| `corrupted_vqav2` | **0.767** | 0.039 | 0.025 | 0.21 |
|
| 117 |
+
| `ood_beamrider` | **0.350** | 0.081 | 0.000 | 0.00 |
|
| 118 |
+
| `ood_boxing` | **0.162** | 0.038 | 0.000 | 0.00 |
|
| 119 |
+
| `ood_cars` | **0.872** | 0.054 | 0.000 | 0.75 |
|
| 120 |
+
| `ood_flowers` | **0.734** | 0.083 | 0.022 | 0.53 |
|
| 121 |
+
| `ood_food101` | **0.886** | 0.010 | 0.012 | 0.81 |
|
| 122 |
+
| `ood_qbert` | **0.278** | 0.011 | 0.000 | 0.00 |
|
| 123 |
+
| `ood_seaquest` | **0.160** | 0.027 | 0.000 | 0.00 |
|
| 124 |
+
| `unknowable` (uniform target) | mean confidence 0.396 vs chance 0.227 | | answered at ≥ 0.9: **1.2%** | |
|
| 125 |
+
|
| 126 |
+
## Limits
|
| 127 |
+
|
| 128 |
+
- **Control does not transfer**: unseen games stay near chance, and on trained games Sev is still below simply repeating
|
| 129 |
+
the previous action (Pong 0.52 vs ~0.69). It is cautious there (no confident errors), not competent.
|
| 130 |
+
- **Degraded satellite images**: on blurred / noised EuroSAT tiles Sev is over-confident (18.7% confident errors).
|
| 131 |
+
- **Aesthetics (AVA) is poorly calibrated** (ECE 0.40): a subjective target.
|
| 132 |
+
- Zero-shot classification of unseen classes is good (cars 0.87, food 0.89, flowers 0.73) but 1-4 points below the
|
| 133 |
+
20k-record Kev start: more task-specific training erodes a little general knowledge.
|
| 134 |
+
- **Hardware**: on Hopper GPUs (H100) flash-linear-attention refuses the Gated DeltaNet backward with torch 2.8's Triton
|
| 135 |
+
(incorrect results, fla #640): train on Ampere/Ada (A100, L4). On Macs the DeltaNet layers run a slow PyTorch fallback.
|
| 136 |
+
- It answers the questions it is asked; it does not verify identity, and it must not be used to infer sensitive
|
| 137 |
+
attributes (the training data asks apparent age only; gender and ethnicity were deliberately never asked).
|
| 138 |
+
|
| 139 |
+
## Licence
|
| 140 |
+
|
| 141 |
+
The adapter and head are derived from Kev-0.8B (Apache-2.0) and Qwen3.5 (Apache-2.0), but were trained on data that
|
| 142 |
+
includes research-only / non-commercial sources (RVL-CDIP, AVA, ScienceQA, AG News, Yelp; see the dataset card).
|
| 143 |
+
**Use for research only** unless you retrain on the public tier (`Jacqkues/kev-vision-decisions`).
|