File size: 8,610 Bytes
d54f613
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
---
license: gemma
base_model: google/gemma-4-E4B-it
library_name: peft
pipeline_tag: text-classification
tags:
  - lora
  - qlora
  - calibration
  - decision-model
  - system-one
  - typesafe
language:
  - en
---

# jev-lite

A QLoRA adapter that turns Gemma 4 E4B into a **System One** decision model: it reads a
state, reads a typed question about it, and returns a calibrated probability distribution
over the allowed answers — from **one forward pass**, with no generation.

Because the answer is read from the logits at the option letters, it **cannot** answer
outside the options it was given. There is no parsing, no retry loop, and no
"as an AI language model".

It implements the three [TypeSafe](https://docs.typesafe.ai) primitives:

| type | question | returns |
|---|---|---|
| `choice` | pick one of several named options | the option, probabilities, confidence |
| `score` | rate against ordered levels | expected level, legend, probabilities, confidence |
| `noul` | is this true? | a single probability |

## Results

Evaluated on 1,898 held-out rows — 31 Super-NaturalInstructions **tasks** the model never
saw, 78 synthetic **states** it never saw, and SST-2 held out in its entirety. The split is
by task and by state, never by row, so these are questions of kinds it was not trained on.

| | accuracy | ECE | NLL | Brier |
|---|---|---|---|---|
| **overall** | **0.816** | **0.019** | 0.430 | 0.222 |
| choice (n=1196) | 0.809 | — | 0.431 | 0.238 |
| noul (n=547) | 0.819 | — | — | — |
| score (n=155) | 0.852 | — | — | — |

Score answers are off by **0.216 levels** on average (mean absolute error of the expected
level).

**Calibration is the point.** ECE of 0.019 means that when it reports 80% confidence it is
right about 80% of the time. Accuracy was flat from step 250 to the end of training while
ECE fell 0.086 → 0.019: the model did not learn to be right more often, it learned to be
honest about when it isn't.

Operationally, gating on confidence:

| band | share of traffic | accuracy |
|---|---|---|
| ≥ 0.80 — act automatically | 62% | 93% |
| 0.50–0.80 — confirm or review | 36% | 60% |
| < 0.50 — route to a human | 2% | 32% |

## Confidence

`confidence` is the model's probability that the answer it returned is the correct one:

- **choice** — the probability of the selected option (`p_max`)
- **score** — the probability mass that rounds to the reported expected level
- **noul** — no confidence field; the probability *is* the answer

This was chosen by measurement, not taste. Over 1,430 gold-labelled held-out rows,
`p_max` beat normalized entropy, margin, chance-corrected top and collision entropy on
**both** AUROC (0.816, ranking right answers above wrong ones) and ECE (0.036). Normalized
entropy — the obvious first guess — was the worst of the lot: it is dominated by small
probabilities, so it reads a decisive 0.85/0.08/0.07 as low confidence and routes 45% of
traffic to human review at 64% accuracy.

## Usage

The adapter is only meaningful with the exact prompt format it was trained on, so
[`primitives.py`](./primitives.py) is included in this repo and **must** be used to render
questions. Criteria descriptions are part of that format.

```python
import torch, primitives
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoTokenizer, BitsAndBytesConfig

tok = AutoTokenizer.from_pretrained("vagmi/jev-lite")
base = AutoModelForImageTextToText.from_pretrained(
    "google/gemma-4-E4B-it", device_map={"": 0}, dtype=torch.bfloat16,
    quantization_config=BitsAndBytesConfig(
        load_in_4bit=True, bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True))
model = PeftModel.from_pretrained(base, "vagmi/jev-lite").eval()

row = primitives.normalize({
    "type": "choice",
    "state": "Help! My payouts have been failing for 3 days.",
    "question": "Which team should handle this?",
    "options": ["billing", "technical", "sales"],
    "criteria": {"billing": "Payments, invoicing, refunds",
                 "technical": "Bugs, outages, integrations",
                 "sales": "Pricing, upgrades, new accounts"},
})

prompt = primitives.PREFIX + row["state"] + "\n</state>\n\n" + \
    primitives.question_block(row) + "\nAnswer:"
ids = torch.tensor([[tok.bos_token_id] + tok.encode(prompt, add_special_tokens=False)])
letters = [tok.encode(" " + c, add_special_tokens=False)[0] for c in primitives.LETTERS]

with torch.no_grad():
    logits = model(input_ids=ids.to(model.device)).logits[0, -1].float()
probs = torch.softmax(logits[letters[:len(row["options"])]], -1).tolist()
print(primitives.answer(row, probs))
# {'type': 'choice', 'choice': 'billing', 'probabilities': {...}, 'confidence': 0.58}
```

A server speaking the TypeSafe System One wire API (`POST /v1/systemone`), for both
transformers and vLLM backends, is at
[github.com/vagmi/jevlite](https://github.com/vagmi/jevlite).

## Training

QLoRA on `google/gemma-4-E4B-it`, 4-bit NF4 base with bf16 compute.

| | |
|---|---|
| LoRA rank / alpha | 16 / 32, dropout 0.05 |
| target modules | q, k, v, o, gate, up, down |
| trainable params | 34.9M of 7.98B (0.44%) |
| objective | soft-label cross-entropy over option letters |
| epochs / LR / schedule | 1 / 2e-4 / cosine, 3% warmup |
| effective batch | 16 (grad accum) |
| hardware | one RTX 4090 (24 GB), ~1h40m |

Options are shuffled every epoch so the model cannot learn "A is usually right" — except
for `score`, where the order of levels carries meaning and is preserved.

## Training data

23,632 rows across 301 tasks: 13,492 choice, 6,368 noul, 3,772 score. 92% carry criteria
descriptions; 35.5% carry soft labels.

| source | rows | labels |
|---|---|---|
| Super-NaturalInstructions (288 tasks) | 8,629 | gold, with criteria written by the teacher |
| MNLI / ANLI / BoolQ / RACE / Yelp | 9,605 | gold, hand-written criteria |
| synthetic states (978, 8 domains) | 5,398 | soft labels from the teacher |

Soft labels came from **Qwen3.6-35B-A3B** read at the answer token with option order
reversed and averaged to cancel position bias. Reasoning was deliberately left **off**: with
a thinking budget the teacher commits to its answer with probability 1.0 on essentially
every row, which is useless for distillation — the measured mean entropy was 0.000 with
reasoning versus 0.211 without.

## Limitations

- **English only.**
- **Score evaluation rests on teacher labels, not gold.** All 155 held-out score rows are
  synthetic; SNI has no rating scales and the HF eval split is SST-2. Choice and noul
  numbers (n=1,430) are against real gold labels.
- **Serving precision matters.** The adapter was trained against a 4-bit NF4 base and
  learned to correct that quantization. Served in bf16 (e.g. vLLM), accuracy holds (0.807
  vs 0.799) but ECE roughly doubles, 0.036 → 0.069. A single temperature of 1.4, fitted on
  half the eval set and measured on the other half, restores it to 0.033. Serve 4-bit, or
  apply the correction.
- **Tokenization is part of the contract.** Feeding a differently-tokenized prompt (a
  missing BOS, say) moved 16% of argmaxes in testing. Build ids the way `primitives.py`
  does.
- **The teacher disagreed with gold on 24.9%** of a 3,000-row sample — mostly noisy
  annotation in SBIC/civil-comments-style tasks plus adversarial ANLI. Those rows were
  blended toward gold at 0.5 rather than dropped.
- Not evaluated for fairness, toxicity, or adversarial robustness. Several source tasks
  (stereotype and offensiveness classification) carry the biases of their annotators.

## Licensing and provenance

The adapter is a derivative of Gemma and is governed by the
[Gemma Terms of Use](https://ai.google.dev/gemma/terms), including its use restrictions.

Training data licenses are **mixed, and some are non-commercial**:

| dataset | license |
|---|---|
| Super-NaturalInstructions | Apache-2.0 |
| ANLI | CC BY-NC 4.0 — **non-commercial** |
| RACE | research use only — **non-commercial** |
| BoolQ | CC BY-SA 3.0 |
| MNLI (GLUE) | mixed, per-genre |
| Yelp Review Full | Yelp Dataset Terms of Use |

Because ANLI and RACE rows are in the training mix, **this adapter should be treated as
non-commercial / research use** unless retrained without them. `build_data.py` in the
source repo takes `--sets` to select sources, so a commercially-clean rebuild is a flag
change plus a retrain.

Synthetic rows and soft labels were generated with Qwen3.6-35B-A3B, an open-weight model
whose license permits using outputs to train other models.