vagmi commited on
Commit
d54f613
·
verified ·
1 Parent(s): d711e75

add README.md

Browse files
Files changed (1) hide show
  1. README.md +200 -0
README.md ADDED
@@ -0,0 +1,200 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: gemma
3
+ base_model: google/gemma-4-E4B-it
4
+ library_name: peft
5
+ pipeline_tag: text-classification
6
+ tags:
7
+ - lora
8
+ - qlora
9
+ - calibration
10
+ - decision-model
11
+ - system-one
12
+ - typesafe
13
+ language:
14
+ - en
15
+ ---
16
+
17
+ # jev-lite
18
+
19
+ A QLoRA adapter that turns Gemma 4 E4B into a **System One** decision model: it reads a
20
+ state, reads a typed question about it, and returns a calibrated probability distribution
21
+ over the allowed answers — from **one forward pass**, with no generation.
22
+
23
+ Because the answer is read from the logits at the option letters, it **cannot** answer
24
+ outside the options it was given. There is no parsing, no retry loop, and no
25
+ "as an AI language model".
26
+
27
+ It implements the three [TypeSafe](https://docs.typesafe.ai) primitives:
28
+
29
+ | type | question | returns |
30
+ |---|---|---|
31
+ | `choice` | pick one of several named options | the option, probabilities, confidence |
32
+ | `score` | rate against ordered levels | expected level, legend, probabilities, confidence |
33
+ | `noul` | is this true? | a single probability |
34
+
35
+ ## Results
36
+
37
+ Evaluated on 1,898 held-out rows — 31 Super-NaturalInstructions **tasks** the model never
38
+ saw, 78 synthetic **states** it never saw, and SST-2 held out in its entirety. The split is
39
+ by task and by state, never by row, so these are questions of kinds it was not trained on.
40
+
41
+ | | accuracy | ECE | NLL | Brier |
42
+ |---|---|---|---|---|
43
+ | **overall** | **0.816** | **0.019** | 0.430 | 0.222 |
44
+ | choice (n=1196) | 0.809 | — | 0.431 | 0.238 |
45
+ | noul (n=547) | 0.819 | — | — | — |
46
+ | score (n=155) | 0.852 | — | — | — |
47
+
48
+ Score answers are off by **0.216 levels** on average (mean absolute error of the expected
49
+ level).
50
+
51
+ **Calibration is the point.** ECE of 0.019 means that when it reports 80% confidence it is
52
+ right about 80% of the time. Accuracy was flat from step 250 to the end of training while
53
+ ECE fell 0.086 → 0.019: the model did not learn to be right more often, it learned to be
54
+ honest about when it isn't.
55
+
56
+ Operationally, gating on confidence:
57
+
58
+ | band | share of traffic | accuracy |
59
+ |---|---|---|
60
+ | ≥ 0.80 — act automatically | 62% | 93% |
61
+ | 0.50–0.80 — confirm or review | 36% | 60% |
62
+ | < 0.50 — route to a human | 2% | 32% |
63
+
64
+ ## Confidence
65
+
66
+ `confidence` is the model's probability that the answer it returned is the correct one:
67
+
68
+ - **choice** — the probability of the selected option (`p_max`)
69
+ - **score** — the probability mass that rounds to the reported expected level
70
+ - **noul** — no confidence field; the probability *is* the answer
71
+
72
+ This was chosen by measurement, not taste. Over 1,430 gold-labelled held-out rows,
73
+ `p_max` beat normalized entropy, margin, chance-corrected top and collision entropy on
74
+ **both** AUROC (0.816, ranking right answers above wrong ones) and ECE (0.036). Normalized
75
+ entropy — the obvious first guess — was the worst of the lot: it is dominated by small
76
+ probabilities, so it reads a decisive 0.85/0.08/0.07 as low confidence and routes 45% of
77
+ traffic to human review at 64% accuracy.
78
+
79
+ ## Usage
80
+
81
+ The adapter is only meaningful with the exact prompt format it was trained on, so
82
+ [`primitives.py`](./primitives.py) is included in this repo and **must** be used to render
83
+ questions. Criteria descriptions are part of that format.
84
+
85
+ ```python
86
+ import torch, primitives
87
+ from peft import PeftModel
88
+ from transformers import AutoModelForImageTextToText, AutoTokenizer, BitsAndBytesConfig
89
+
90
+ tok = AutoTokenizer.from_pretrained("vagmi/jev-lite")
91
+ base = AutoModelForImageTextToText.from_pretrained(
92
+ "google/gemma-4-E4B-it", device_map={"": 0}, dtype=torch.bfloat16,
93
+ quantization_config=BitsAndBytesConfig(
94
+ load_in_4bit=True, bnb_4bit_quant_type="nf4",
95
+ bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True))
96
+ model = PeftModel.from_pretrained(base, "vagmi/jev-lite").eval()
97
+
98
+ row = primitives.normalize({
99
+ "type": "choice",
100
+ "state": "Help! My payouts have been failing for 3 days.",
101
+ "question": "Which team should handle this?",
102
+ "options": ["billing", "technical", "sales"],
103
+ "criteria": {"billing": "Payments, invoicing, refunds",
104
+ "technical": "Bugs, outages, integrations",
105
+ "sales": "Pricing, upgrades, new accounts"},
106
+ })
107
+
108
+ prompt = primitives.PREFIX + row["state"] + "\n</state>\n\n" + \
109
+ primitives.question_block(row) + "\nAnswer:"
110
+ ids = torch.tensor([[tok.bos_token_id] + tok.encode(prompt, add_special_tokens=False)])
111
+ letters = [tok.encode(" " + c, add_special_tokens=False)[0] for c in primitives.LETTERS]
112
+
113
+ with torch.no_grad():
114
+ logits = model(input_ids=ids.to(model.device)).logits[0, -1].float()
115
+ probs = torch.softmax(logits[letters[:len(row["options"])]], -1).tolist()
116
+ print(primitives.answer(row, probs))
117
+ # {'type': 'choice', 'choice': 'billing', 'probabilities': {...}, 'confidence': 0.58}
118
+ ```
119
+
120
+ A server speaking the TypeSafe System One wire API (`POST /v1/systemone`), for both
121
+ transformers and vLLM backends, is at
122
+ [github.com/vagmi/jevlite](https://github.com/vagmi/jevlite).
123
+
124
+ ## Training
125
+
126
+ QLoRA on `google/gemma-4-E4B-it`, 4-bit NF4 base with bf16 compute.
127
+
128
+ | | |
129
+ |---|---|
130
+ | LoRA rank / alpha | 16 / 32, dropout 0.05 |
131
+ | target modules | q, k, v, o, gate, up, down |
132
+ | trainable params | 34.9M of 7.98B (0.44%) |
133
+ | objective | soft-label cross-entropy over option letters |
134
+ | epochs / LR / schedule | 1 / 2e-4 / cosine, 3% warmup |
135
+ | effective batch | 16 (grad accum) |
136
+ | hardware | one RTX 4090 (24 GB), ~1h40m |
137
+
138
+ Options are shuffled every epoch so the model cannot learn "A is usually right" — except
139
+ for `score`, where the order of levels carries meaning and is preserved.
140
+
141
+ ## Training data
142
+
143
+ 23,632 rows across 301 tasks: 13,492 choice, 6,368 noul, 3,772 score. 92% carry criteria
144
+ descriptions; 35.5% carry soft labels.
145
+
146
+ | source | rows | labels |
147
+ |---|---|---|
148
+ | Super-NaturalInstructions (288 tasks) | 8,629 | gold, with criteria written by the teacher |
149
+ | MNLI / ANLI / BoolQ / RACE / Yelp | 9,605 | gold, hand-written criteria |
150
+ | synthetic states (978, 8 domains) | 5,398 | soft labels from the teacher |
151
+
152
+ Soft labels came from **Qwen3.6-35B-A3B** read at the answer token with option order
153
+ reversed and averaged to cancel position bias. Reasoning was deliberately left **off**: with
154
+ a thinking budget the teacher commits to its answer with probability 1.0 on essentially
155
+ every row, which is useless for distillation — the measured mean entropy was 0.000 with
156
+ reasoning versus 0.211 without.
157
+
158
+ ## Limitations
159
+
160
+ - **English only.**
161
+ - **Score evaluation rests on teacher labels, not gold.** All 155 held-out score rows are
162
+ synthetic; SNI has no rating scales and the HF eval split is SST-2. Choice and noul
163
+ numbers (n=1,430) are against real gold labels.
164
+ - **Serving precision matters.** The adapter was trained against a 4-bit NF4 base and
165
+ learned to correct that quantization. Served in bf16 (e.g. vLLM), accuracy holds (0.807
166
+ vs 0.799) but ECE roughly doubles, 0.036 → 0.069. A single temperature of 1.4, fitted on
167
+ half the eval set and measured on the other half, restores it to 0.033. Serve 4-bit, or
168
+ apply the correction.
169
+ - **Tokenization is part of the contract.** Feeding a differently-tokenized prompt (a
170
+ missing BOS, say) moved 16% of argmaxes in testing. Build ids the way `primitives.py`
171
+ does.
172
+ - **The teacher disagreed with gold on 24.9%** of a 3,000-row sample — mostly noisy
173
+ annotation in SBIC/civil-comments-style tasks plus adversarial ANLI. Those rows were
174
+ blended toward gold at 0.5 rather than dropped.
175
+ - Not evaluated for fairness, toxicity, or adversarial robustness. Several source tasks
176
+ (stereotype and offensiveness classification) carry the biases of their annotators.
177
+
178
+ ## Licensing and provenance
179
+
180
+ The adapter is a derivative of Gemma and is governed by the
181
+ [Gemma Terms of Use](https://ai.google.dev/gemma/terms), including its use restrictions.
182
+
183
+ Training data licenses are **mixed, and some are non-commercial**:
184
+
185
+ | dataset | license |
186
+ |---|---|
187
+ | Super-NaturalInstructions | Apache-2.0 |
188
+ | ANLI | CC BY-NC 4.0 — **non-commercial** |
189
+ | RACE | research use only — **non-commercial** |
190
+ | BoolQ | CC BY-SA 3.0 |
191
+ | MNLI (GLUE) | mixed, per-genre |
192
+ | Yelp Review Full | Yelp Dataset Terms of Use |
193
+
194
+ Because ANLI and RACE rows are in the training mix, **this adapter should be treated as
195
+ non-commercial / research use** unless retrained without them. `build_data.py` in the
196
+ source repo takes `--sets` to select sources, so a commercially-clean rebuild is a flag
197
+ change plus a retrain.
198
+
199
+ Synthetic rows and soft labels were generated with Qwen3.6-35B-A3B, an open-weight model
200
+ whose license permits using outputs to train other models.