Qwen3.5-4B, readout fine-tuned (full, + coherence, lr 1e-6 as selected at 2B)
A System One decision model: it reads a state, answers typed questions (choice, score, noul) and returns calibrated
probability distributions your code can branch on — it never writes text. This repo is every parameter of Qwen/Qwen3.5-4B fine-tuned, trained on its own decision
readout with a coherence penalty.
At a glance — accuracy 0.757 · ECE 0.056 · held-out 0.802 · TVD to human labels 0.303 · sure loss 0.024
same order, Qwen3.5-4B, untuned (Tier 0): 0.662 / 0.090 / 0.719 / 0.438 / 0.151
same order, Jev 1.13.0: 0.733 / 0.113 / 0.835 / 0.432 / 0.081
jev-bench · leaderboard · findings · code
Use it
from jevify import load_jevified
model = load_jevified("Praveenrajus/jevify-qwen3.5-4b-readout-full-coh")
model.ask({"text": "The battery lasted two days on a single charge."},
{"q": {"type": "noul", "instructions": "Is the review positive?"}})
jevify-serve --model Praveenrajus/jevify-qwen3.5-4b-readout-full-coh serves it as a drop-in for the TypeSafe SDK (TYPESAFE_BASE_URL=http://localhost:8000).
The weights are in this repo.
Results
Every number is on the jev-bench test splits (22,773 records) or the study's other test suites, scored the same way for every model; the rows under this model are references from the same study.
Decisions and calibration
| model | acc | ECE | Brier | held-out acc | TVD to human labels |
|---|---|---|---|---|---|
| this model | 0.757 | 0.056 | 0.310 | 0.802 | 0.303 |
| Qwen3.5-4B, untuned (Tier 0) | 0.662 | 0.090 | 0.401 | 0.719 | 0.438 |
| full fine-tune, supervised only (same lr) | 0.752 | 0.061 | 0.318 | 0.794 | 0.322 |
| LoRA + coherence, same data and seed | 0.751 | 0.058 | 0.316 | 0.792 | 0.303 |
| Jev 1.13.0 (TypeSafe API) | 0.733 | 0.113 | 0.349 | 0.835 | 0.432 |
Coherence and invariance — sure loss: mean d² over 4,749 question families (0 = perfectly coherent); order flip: how often the top answer changes when options are shuffled; tag TVD: how much the distribution moves when option tags change from A–J to other identifiers.
| model | sure loss | share incoherent | order flip | tag TVD | K=2→max acc drop |
|---|---|---|---|---|---|
| this model | 0.024 | 0.328 | 0.064 | 0.022 | 0.230 |
| Qwen3.5-4B, untuned (Tier 0) | 0.151 | 0.959 | 0.140 | 0.038 | 0.299 |
| full fine-tune, supervised only (same lr) | 0.316 | 0.913 | 0.068 | 0.024 | 0.235 |
| LoRA + coherence, same data and seed | 0.029 | 0.347 | 0.057 | 0.023 | 0.230 |
| Jev 1.13.0 (TypeSafe API) | 0.081 | 0.725 | 0.046 | — | 0.246 |
Out of distribution — stated rules (LegalBench, rule given in the question), none-of-the-above when the gold option is removed, injected-instruction hijack rate, and three community Jev benchmarks.
| model | stated rule | 'none' when gone | hijack | phishing AUROC | tool risk |
|---|---|---|---|---|---|
| this model | 0.656 | 0.604 | 0.057 | 0.815 | 0.900 |
| Qwen3.5-4B, untuned (Tier 0) | 0.619 | 0.484 | 0.394 | 0.784 | 0.867 |
| full fine-tune, supervised only (same lr) | 0.759 | 0.560 | 0.054 | 0.836 | 0.883 |
| LoRA + coherence, same data and seed | 0.723 | 0.682 | 0.089 | 0.929 | 0.900 |
| Jev 1.13.0 (TypeSafe API) | 0.924 | 0.744 | 0.205 | 0.688 | 0.933 |
Reproduction check. Loading this folder with load_jevified and re-scoring 72 jev-bench test records from six sources reproduced the training run's own test predictions: 0 changed choice answers, mean largest |Δp| 0.007, max 0.031.
How it was trained
The model is trained on its own decision readout — the distribution over the allowed answers read at the answer position,
one forward pass, no decoding — with the primitive's proper scoring rule, plus a coherence penalty (weight 1.0): every training question comes with automatically derived siblings (the options as yes/no questions, the negation, the threshold questions of a scale), and the de Finetti sure loss of the family's answers is penalised, so the model's answers to related questions stay mutually consistent.
Options are shuffled per family. Training data: the train splits of the 16 non-held-out jev-bench sources
(5,885 families, at most 400 records per source); lr 1e-06, 2 epochs,
best epoch by validation loss (epoch 0), seed 0, fp32 master weights with bf16 autocast.
A Tier 0 recipe (temperature per primitive, Noul bias, option-order permutations) was then fitted on validation splits.
The six held-out sources (clinc150, arc_challenge, yelp5, measuring_hate_speech, fever_evidence, strategyqa_grounded) never appeared in training.
Files
jevify_config.json— the recipe, the backbone and the training settingsload_jevifiedreads- model weights, tokenizer and chat template — the full fine-tuned checkpoint
results/test_metrics.json— every jev-bench config;recipe.json— the fitted reciperesults/coherence.json,probes.json,tags.json— the coherence, probe and tag testsresults/train.json— the training log;summary.json— this model's row of the study tableresults/verification.json— the reproduction check reported under Results
Related models
Limitations
- One training seed per repo branch; out-of-distribution numbers in particular vary between identical runs, so compare arms across seeds before drawing conclusions.
- The phishing benchmark's decision threshold shifts after fine-tuning (ranking, AUROC, is preserved); a one-number log-odds shift fitted on a handful of labelled emails repairs it.
- English only; the recipe was fitted on jev-bench validation splits and may need refitting on a very different domain.
- Downloads last month
- 20