File size: 6,212 Bytes
553e496
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
---
license: apache-2.0
base_model: HuggingFaceTB/SmolLM3-3B-checkpoints
library_name: jevify
tags: [jevify, system-one, decision-model, calibration, coherence]
datasets: [Praveenrajus/jev-bench]
---

# SmolLM3-3B (APO checkpoint), readout fine-tuned (LoRA, supervised)

A **System One decision model**: it reads a `state`, answers typed questions (`choice`, `score`, `noul`) and returns calibrated
probability distributions your code can branch on β€” it never writes text. This repo is a rank-16 LoRA (30,228,480 parameters) on `HuggingFaceTB/SmolLM3-3B-checkpoints`, merged into the weights at load, trained on its own decision
readout.

> **At a glance** β€” accuracy **0.705** Β· ECE **0.058** Β· held-out **0.741** Β· TVD to human labels **0.337** Β· sure loss **0.292**
> <br>same order, SmolLM3-3B APO checkpoint, untuned (Tier 0): 0.539 / 0.117 / 0.579 / 0.460 / 0.226<br>same order, Jev 1.13.0: 0.733 / 0.113 / 0.835 / 0.432 / 0.081

[jev-bench](https://huggingface.co/datasets/Praveenrajus/jev-bench) Β· [leaderboard](https://huggingface.co/datasets/Praveenrajus/jev-bench#leaderboard) Β· [findings](https://github.com/uspraveen/Jevify/blob/main/docs/FINDINGS.md#18-readout-fine-tuning-and-what-a-coherence-penalty-adds) Β· [code](https://github.com/uspraveen/Jevify)

## Use it

```python
from jevify import load_jevified

model = load_jevified("Praveenrajus/jevify-smollm3-3b-apo-readout")
model.ask({"text": "The battery lasted two days on a single charge."},
          {"q": {"type": "noul", "instructions": "Is the review positive?"}})
```

`jevify-serve --model Praveenrajus/jevify-smollm3-3b-apo-readout` serves it as a drop-in for the TypeSafe SDK (`TYPESAFE_BASE_URL=http://localhost:8000`).
The backbone is pulled from its own repo at load, pinned to commit `cfb32d505f5025ec9be4e704f70cfbf5bdf8da94`.

## Results

Every number is on the jev-bench **test** splits (22,773 records) or the study's other test suites, scored the same way for
every model; the rows under this model are references from the same study.

**Decisions and calibration**

| model | acc | ECE | Brier | held-out acc | TVD to human labels |
|---|---|---|---|---|---|
| **this model** | 0.705 | 0.058 | 0.363 | 0.741 | 0.337 |
| SmolLM3-3B APO checkpoint, untuned (Tier 0) | 0.539 | 0.117 | 0.526 | 0.579 | 0.460 |
| same recipe + coherence | 0.708 | 0.055 | 0.360 | 0.749 | 0.318 |
| same recipe from the SFT checkpoint | 0.705 | 0.057 | 0.365 | 0.738 | 0.350 |
| Jev 1.13.0 (TypeSafe API) | 0.733 | 0.113 | 0.349 | 0.835 | 0.432 |

**Coherence and invariance** β€” sure loss: mean dΒ² over 4,749 question families (0 = perfectly coherent); order flip: how
often the top answer changes when options are shuffled; tag TVD: how much the distribution moves when option tags change
from A–J to other identifiers.

| model | sure loss | share incoherent | order flip | tag TVD | K=2β†’max acc drop |
|---|---|---|---|---|---|
| **this model** | 0.292 | 0.964 | β€” | β€” | β€” |
| SmolLM3-3B APO checkpoint, untuned (Tier 0) | 0.226 | 0.990 | 0.435 | 0.077 | 0.566 |
| same recipe + coherence | 0.041 | 0.640 | β€” | β€” | β€” |
| same recipe from the SFT checkpoint | 0.315 | 0.966 | β€” | β€” | β€” |
| Jev 1.13.0 (TypeSafe API) | 0.081 | 0.725 | 0.046 | β€” | 0.246 |

**Out of distribution** β€” stated rules (LegalBench, rule given in the question), none-of-the-above when the gold option is
removed, injected-instruction hijack rate, and three community Jev benchmarks.

| model | stated rule | 'none' when gone | hijack | phishing AUROC | tool risk |
|---|---|---|---|---|---|
| **this model** | β€” | β€” | β€” | β€” | β€” |
| SmolLM3-3B APO checkpoint, untuned (Tier 0) | 0.588 | 0.552 | 0.464 | 0.794 | 0.800 |
| same recipe + coherence | β€” | β€” | β€” | β€” | β€” |
| same recipe from the SFT checkpoint | β€” | β€” | β€” | β€” | β€” |
| Jev 1.13.0 (TypeSafe API) | 0.924 | 0.744 | 0.205 | 0.688 | 0.933 |

**Reproduction check.** Loading this folder with `load_jevified` and re-scoring 72 jev-bench test records from six sources reproduced the training run's own test predictions: 0 changed choice answers, mean largest |Ξ”p| 0.004, max 0.030 (the adapter is merged into bf16 weights at load).

## How it was trained

The model is trained on its own *decision readout* β€” the distribution over the allowed answers read at the answer position,
one forward pass, no decoding β€” with the primitive's proper scoring rule.
Options are shuffled per family. Training data: the train splits of the 16 non-held-out jev-bench sources
(5,885 families, at most 400 records per source); lr 3e-05, 2 epochs,
best epoch by validation loss (epoch 1), seed 0.
A Tier 0 recipe (temperature per primitive, Noul bias, option-order permutations) was then fitted on validation splits.
The six held-out sources (`clinc150`, `arc_challenge`, `yelp5`, `measuring_hate_speech`, `fever_evidence`, `strategyqa_grounded`) never appeared in training.

## Files

- `jevify_config.json` β€” the recipe, the backbone and the training settings `load_jevified` reads
- `lora/` β€” the adapter, merged into the backbone at load
- `results/test_metrics.json` β€” every jev-bench config; `recipe.json` β€” the fitted recipe
- `results/coherence.json`, `probes.json`, `tags.json` β€” the coherence, probe and tag tests
- `results/train.json` β€” the training log; `summary.json` β€” this model's row of the study table
- `results/verification.json` β€” the reproduction check reported under Results

## Related models

- [Same recipe + coherence penalty](https://huggingface.co/Praveenrajus/jevify-smollm3-3b-apo-readout-coh)
- [Same recipe from the SFT checkpoint (the repair comparison)](https://huggingface.co/Praveenrajus/jevify-smollm3-3b-sft-readout)

## Limitations

- One training seed per repo branch; out-of-distribution numbers in particular vary between identical runs, so compare
  arms across seeds before drawing conclusions.
- The phishing benchmark's decision threshold shifts after fine-tuning (ranking, AUROC, is preserved); a one-number
  log-odds shift fitted on a handful of labelled emails repairs it.
- English only; the recipe was fitted on jev-bench validation splits and may need refitting on a very different domain.