--- license: apache-2.0 base_model: Qwen/Qwen3.5-2B library_name: jevify tags: [jevify, system-one, decision-model, calibration, coherence] datasets: [Praveenrajus/jev-bench] --- # Qwen3.5-2B, readout fine-tuned (full, supervised, lr 1e-6 selected on validation) A **System One decision model**: it reads a `state`, answers typed questions (`choice`, `score`, `noul`) and returns calibrated probability distributions your code can branch on — it never writes text. This repo is every parameter of `Qwen/Qwen3.5-2B` fine-tuned, trained on its own decision readout. > **At a glance** — accuracy **0.703** · ECE **0.056** · held-out **0.746** · TVD to human labels **0.315** · sure loss **0.344** >
same order, Qwen3.5-2B, untuned (Tier 0): 0.577 / 0.089 / 0.634 / 0.458 / 0.212
same order, Jev 1.13.0: 0.733 / 0.113 / 0.835 / 0.432 / 0.081 [jev-bench](https://huggingface.co/datasets/Praveenrajus/jev-bench) · [leaderboard](https://huggingface.co/datasets/Praveenrajus/jev-bench#leaderboard) · [findings](https://github.com/uspraveen/Jevify/blob/main/docs/FINDINGS.md#18-readout-fine-tuning-and-what-a-coherence-penalty-adds) · [code](https://github.com/uspraveen/Jevify) ## Use it ```python from jevify import load_jevified model = load_jevified("Praveenrajus/jevify-qwen3.5-2b-readout-full") model.ask({"text": "The battery lasted two days on a single charge."}, {"q": {"type": "noul", "instructions": "Is the review positive?"}}) ``` `jevify-serve --model Praveenrajus/jevify-qwen3.5-2b-readout-full` serves it as a drop-in for the TypeSafe SDK (`TYPESAFE_BASE_URL=http://localhost:8000`). The weights are in this repo. ## Results Every number is on the jev-bench **test** splits (22,773 records) or the study's other test suites, scored the same way for every model; the rows under this model are references from the same study. **Decisions and calibration** | model | acc | ECE | Brier | held-out acc | TVD to human labels | |---|---|---|---|---|---| | **this model** | 0.703 | 0.056 | 0.358 | 0.746 | 0.315 | | Qwen3.5-2B, untuned (Tier 0) | 0.577 | 0.089 | 0.485 | 0.634 | 0.458 | | LoRA, same data and seed | 0.701 | 0.051 | 0.362 | 0.736 | 0.324 | | full fine-tune + coherence (same lr) | 0.704 | 0.056 | 0.358 | 0.742 | 0.307 | | Jev 1.13.0 (TypeSafe API) | 0.733 | 0.113 | 0.349 | 0.835 | 0.432 | **Coherence and invariance** — sure loss: mean d² over 4,749 question families (0 = perfectly coherent); order flip: how often the top answer changes when options are shuffled; tag TVD: how much the distribution moves when option tags change from A–J to other identifiers. | model | sure loss | share incoherent | order flip | tag TVD | K=2→max acc drop | |---|---|---|---|---|---| | **this model** | 0.344 | 0.980 | 0.113 | 0.050 | 0.253 | | Qwen3.5-2B, untuned (Tier 0) | 0.212 | 0.980 | 0.271 | 0.048 | 0.452 | | LoRA, same data and seed | 0.339 | 0.972 | 0.104 | 0.044 | 0.259 | | full fine-tune + coherence (same lr) | 0.028 | 0.486 | 0.109 | 0.049 | 0.250 | | Jev 1.13.0 (TypeSafe API) | 0.081 | 0.725 | 0.046 | — | 0.246 | **Out of distribution** — stated rules (LegalBench, rule given in the question), none-of-the-above when the gold option is removed, injected-instruction hijack rate, and three community Jev benchmarks. | model | stated rule | 'none' when gone | hijack | phishing AUROC | tool risk | |---|---|---|---|---|---| | **this model** | 0.671 | 0.336 | 0.055 | 0.775 | 0.783 | | Qwen3.5-2B, untuned (Tier 0) | 0.602 | 0.850 | 0.269 | 0.902 | 0.733 | | LoRA, same data and seed | 0.536 | 0.472 | 0.055 | 0.776 | 0.767 | | full fine-tune + coherence (same lr) | 0.612 | 0.462 | 0.021 | 0.772 | 0.800 | | Jev 1.13.0 (TypeSafe API) | 0.924 | 0.744 | 0.205 | 0.688 | 0.933 | **Reproduction check.** Loading this folder with `load_jevified` and re-scoring 72 jev-bench test records from six sources reproduced the training run's own test predictions: 0 changed choice answers, mean largest |Δp| 0.007, max 0.038. ## How it was trained The model is trained on its own *decision readout* — the distribution over the allowed answers read at the answer position, one forward pass, no decoding — with the primitive's proper scoring rule. Options are shuffled per family. Training data: the train splits of the 16 non-held-out jev-bench sources (5,885 families, at most 400 records per source); lr 1e-06, 2 epochs, best epoch by validation loss (epoch 0), seed 0, fp32 master weights with bf16 autocast. A Tier 0 recipe (temperature per primitive, Noul bias, option-order permutations) was then fitted on validation splits. The six held-out sources (`clinc150`, `arc_challenge`, `yelp5`, `measuring_hate_speech`, `fever_evidence`, `strategyqa_grounded`) never appeared in training. ## Files - `jevify_config.json` — the recipe, the backbone and the training settings `load_jevified` reads - model weights, tokenizer and chat template — the full fine-tuned checkpoint - `results/test_metrics.json` — every jev-bench config; `recipe.json` — the fitted recipe - `results/coherence.json`, `probes.json`, `tags.json` — the coherence, probe and tag tests - `results/train.json` — the training log; `summary.json` — this model's row of the study table - `results/verification.json` — the reproduction check reported under Results ## Related models - [Same recipe + coherence penalty](https://huggingface.co/Praveenrajus/jevify-qwen3.5-2b-readout-full-coh) - [LoRA instead of full fine-tuning](https://huggingface.co/Praveenrajus/jevify-qwen3.5-2b-readout ) ## Limitations - One training seed per repo branch; out-of-distribution numbers in particular vary between identical runs, so compare arms across seeds before drawing conclusions. - The phishing benchmark's decision threshold shifts after fine-tuning (ranking, AUROC, is preserved); a one-number log-odds shift fitted on a handful of labelled emails repairs it. - English only; the recipe was fitted on jev-bench validation splits and may need refitting on a very different domain.