Visual Question Answering
PEFT
Safetensors
English
decision-model
calibration
typesafe
kev
jev
vision
qwen3.5

Sev-0.8B

Sev is a decision model that reads images: one image plus a state (text or JSON) and a set of typed questions go in, a calibrated probability distribution per question comes out, in a single forward pass, with no text generation. Question types are the TypeSafe /v1/systemone ones: choice (one of named options), noul (yes / no) and score (ordered levels). New questions and options are given at request time; nothing is retrained.

It is Kev-0.8B (a Jev-style text decision model: LoRA + pointer head on Qwen/Qwen3.5-0.8B-Base) plugged into the full Qwen3.5 vision-language model and fine-tuned on kev-vision-decisions-full (gated: research use, per-source licences). The vision tower is frozen; the Kev LoRA (r = 16) and pointer head are trained further.

  • Status: a research project, heavily vibe-coded with Claude Code. Not for production.
  • Document version: Sev-0.8B-docindex (page indexing: 21 attributes)
  • Code: kev-vision (VisionDecisionModel, training, dataset builder, evaluation)
  • Served temperature: T = 1.32, fitted on the validation split (hard-label questions only)

Use

from kev.api import SystemOneRequest, to_record
from kev_vision.model import VisionDecisionModel
from PIL import Image

model = VisionDecisionModel("Jacqkues/sev-0.8b", device="cuda").eval()     # Ampere/Ada GPU, MPS or CPU (see limits)
request = {"state": "A document page is attached.",
           "questions": {"doc_type": {"type": "choice", "instructions": "What type of document is this?",
                                      "criteria": {"invoice": "An invoice", "receipt": "A receipt", "letter": "A letter"}},
                         "handwritten": {"type": "noul", "instructions": "Is it mostly handwritten?"}}}
rec, meta = to_record(SystemOneRequest.model_validate(request)); rec["image"] = Image.open("page.jpg")
for m, p in zip(meta, model.probs(rec)):
    print(m["id"], dict(zip(m["keys"], p.tolist())))

How it works

[state] <vision_start> image tokens <vision_end> state text  |  [q] instructions [opt] option ... [decide]
Qwen3.5 (vision frozen, LoRA on the language model)          →  pointer head: <decide> · each option  →  softmax / T

Every question is its own causal row over the shared state (Qwen3.5 is a Gated DeltaNet hybrid), positions are Qwen's 3D M-RoPE, and each option is scored at its end token. On text-only input Sev's code reproduces Kev exactly (max |Δp| 5.7e-7).

Training

  • Start: jaredpalmer/kev-0.8b (adapter + head). A controlled comparison on 20k records picked it over a bare Qwen3.5-0.8B-Base with a fresh LoRA of the same rank: better on 14 of 21 test sets, clearly on text (0.827 vs 0.766).
  • Data: kev-vision-decisions-full, ~126k training records: VQAv2 yes/no + counting, A-OKVQA, ScienceQA, EuroSAT, Oxford Pets, FairFace age, AVA aesthetics (soft targets), ChartNet, document types (DocLayNet, CORD receipts, invoices, RVL-CDIP), control (Pong, Breakout, Freeway, Space Invaders, Ms. Pac-Man from JAT; LeRobot PushT, upweighted ×2), text replay from Kev's sources, and blank / noise images with a uniform target. Yes/no balanced per source, option order randomised, images shared between train and held-out splits removed.
  • Recipe: 1 epoch (136,502 passages), batch 8, AdamW lr 5e-5 OneCycle, bf16 autocast, Kev's augmentation (option permutation, none-of-the-above, distractors) redrawn each pass. ~3 h on one A100 80 GB (Modal).
  • Validation: loss 0.6635, accuracy 0.763, Brier 0.31.

Results: test splits (200 records each, served temperature)

test set before (Kev-0.8B, zero-shot) Sev-0.8B Brier ECE
ai2d_ood 0.585 0.665 0.446 0.077
aokvqa 0.77 0.810 0.267 0.028
ava 0.23 0.640 0.706 0.397
breakout 0.37 0.445 0.650 0.068
chartnet 0.825 0.971 0.043 0.016
countbench_ood 0.555 0.730 0.399 0.106
documents 0.866 0.985 0.025 0.012
eurosat 0.588 0.940 0.087 0.029
fairface_age 0.295 0.595 0.544 0.059
freeway 0.862 0.814 0.359 0.176
mspacman 0.05 0.380 0.768 0.062
pets 0.815 0.975 0.036 0.025
pong 0.462 0.519 0.548 0.048
pope_ood 0.855 0.880 0.178 0.051
pusht 0.213 0.647 0.485 0.084
rvl_cdip 0.567 0.861 0.202 0.031
scienceqa 0.619 0.942 0.102 0.051
spaceinvaders 0.35 0.490 0.667 0.086
text 0.805 0.830 0.243 0.057
text_full 0.844 0.851 0.214 0.053
vqav2 0.796 0.853 0.210 0.042

Blank images (the image carries no answer): mean confidence 0.245 against a uniform 0.238.

Results: sev-ood (never trained on)

Unseen Atari games, unseen classes, questions the image cannot answer, and corrupted images (Jacqkues/sev-ood, gated, 500 records per set). Confident errors = wrong at p ≥ 0.9; coverage = share of decisions that could be automated with at most 5% of them wrong.

set accuracy ECE confident errors coverage at 5% error
corrupted_eurosat 0.555 0.263 0.187 0.00
corrupted_pets 0.938 0.018 0.005 0.98
corrupted_vqav2 0.767 0.039 0.025 0.21
ood_beamrider 0.350 0.081 0.000 0.00
ood_boxing 0.162 0.038 0.000 0.00
ood_cars 0.872 0.054 0.000 0.75
ood_flowers 0.734 0.083 0.022 0.53
ood_food101 0.886 0.010 0.012 0.81
ood_qbert 0.278 0.011 0.000 0.00
ood_seaquest 0.160 0.027 0.000 0.00
unknowable (uniform target) mean confidence 0.396 vs chance 0.227 answered at ≥ 0.9: 1.2%

Results: generic vision benchmarks (never trained on)

Five public benchmarks, cast as typed questions: MMStar (A-D, vision-indispensable), MMBench-en dev (1,000 sampled), MME (yes/no), RealWorldQA (A-F or yes/no) and HallusionBench (image split, yes/no). ScienceQA-sourced items were dropped (Sev trained on ScienceQA). Contamination control: every benchmark image was hashed (256-bit average hash) against every training image; the 305 of 6,001 items within 6 bits of a training image (mostly COCO photos via VQAv2 / A-OKVQA, and ScienceQA) are removed below. Removing them changes no conclusion: they were slightly easier for all three models, including the baseline that never saw them.

The baseline is Qwen3.5-0.8B (the post-trained VLM of the same family) in zero-shot, scored by its next-token probability over the option letters (or Yes / No), so the same calibration metrics apply. All models see images capped at 448 x 448 pixels. Accuracy on the 5,696 clean items; gap = Sev minus Qwen, paired bootstrap 95% interval.

benchmark (chance) n Qwen3.5-0.8B Sev-0.8B gap [95% CI] ECE Qwen → Sev
MMStar (0.25) 1,036 0.460 0.526 +6.6 [+3.3, +9.9] 0.18 → 0.08
MMBench dev (0.26) 927 0.774 0.825 +5.1 [+2.6, +7.6] 0.02 → 0.06
MME (0.50) 2,206 0.711 0.771 +6.0 [+4.2, +7.9] 0.12 → 0.01
RealWorldQA (0.39) 586 0.570 0.631 +6.1 [+2.2, +10.2] 0.15 → 0.05
HallusionBench (0.50) 941 0.644 0.629 −1.5 [−4.8, +1.9] 0.15 → 0.10
all 5,696 0.650 0.698 +4.7 [+3.5, +6.0] 0.12 → 0.02

Confident errors (wrong at p ≥ 0.9), all clean items: 4.5% for Qwen, 1.5% for Sev. HallusionBench (visual illusions, edited charts) is the exception: no model of this size is reliably better than chance-level guessing there.

Results: docbench (document pages, never trained on)

DocLayNet test + validation pages, three yes/no questions (table of contents? figure? table?), 676 items, labels from independent rules. See Sev-0.8B-docindex for the document model.

model accuracy AUROC ECE
Kev-0.8B 0.886 0.934 0.236
Sev-0.8B 0.831 0.924 0.044
Sev-0.8B-docindex 0.928 0.976 0.018

Limits

  • Control does not transfer: unseen games stay near chance, and on trained games Sev is still below simply repeating the previous action (Pong 0.52 vs ~0.69). It is cautious there (no confident errors), not competent.
  • Degraded satellite images: on blurred / noised EuroSAT tiles Sev is over-confident (18.7% confident errors).
  • Aesthetics (AVA) is poorly calibrated (ECE 0.40): a subjective target.
  • Zero-shot classification of unseen classes is good (cars 0.87, food 0.89, flowers 0.73) but 1-4 points below the 20k-record Kev start: more task-specific training erodes a little general knowledge.
  • Hardware: on Hopper GPUs (H100) flash-linear-attention refuses the Gated DeltaNet backward with torch 2.8's Triton (incorrect results, fla #640): train on Ampere/Ada (A100, L4). On Macs the DeltaNet layers run a slow PyTorch fallback.
  • It answers the questions it is asked; it does not verify identity, and it must not be used to infer sensitive attributes (the training data asks apparent age only; gender and ethnicity were deliberately never asked).

Licence

The adapter and head are derived from Kev-0.8B (Apache-2.0) and Qwen3.5 (Apache-2.0), but were trained on data that includes research-only / non-commercial sources (RVL-CDIP, AVA, ScienceQA, AG News, Yelp; see the dataset card). Use for research only. A commercially usable version would have to be retrained on permissively licensed sources only.

Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jacqkues/sev-0.8b

Adapter
(3)
this model

Datasets used to train Jacqkues/sev-0.8b