vagmi commited on
Commit
90b23bc
·
verified ·
1 Parent(s): 5b738f8

Update model card to improve readability . move usage up front.

Browse files
Files changed (1) hide show
  1. README.md +47 -46
README.md CHANGED
@@ -18,9 +18,9 @@ language:
18
 
19
  A QLoRA adapter that turns Gemma 4 E4B into a **System One** decision model: it reads a
20
  state, reads a typed question about it, and returns a calibrated probability distribution
21
- over the allowed answers — from **one forward pass**, with no generation.
22
 
23
- Because the answer is read from the logits at the option letters, it **cannot** answer
24
  outside the options it was given. There is no parsing, no retry loop, and no
25
  "as an AI language model".
26
 
@@ -32,50 +32,6 @@ It implements the three [TypeSafe](https://docs.typesafe.ai) primitives:
32
  | `score` | rate against ordered levels | expected level, legend, probabilities, confidence |
33
  | `noul` | is this true? | a single probability |
34
 
35
- ## Results
36
-
37
- Evaluated on 1,898 held-out rows — 31 Super-NaturalInstructions **tasks** the model never
38
- saw, 78 synthetic **states** it never saw, and SST-2 held out in its entirety. The split is
39
- by task and by state, never by row, so these are questions of kinds it was not trained on.
40
-
41
- | | accuracy | ECE | NLL | Brier |
42
- |---|---|---|---|---|
43
- | **overall** | **0.816** | **0.019** | 0.430 | 0.222 |
44
- | choice (n=1196) | 0.809 | — | 0.431 | 0.238 |
45
- | noul (n=547) | 0.819 | — | — | — |
46
- | score (n=155) | 0.852 | — | — | — |
47
-
48
- Score answers are off by **0.216 levels** on average (mean absolute error of the expected
49
- level).
50
-
51
- **Calibration is the point.** ECE of 0.019 means that when it reports 80% confidence it is
52
- right about 80% of the time. Accuracy was flat from step 250 to the end of training while
53
- ECE fell 0.086 → 0.019: the model did not learn to be right more often, it learned to be
54
- honest about when it isn't.
55
-
56
- Operationally, gating on confidence:
57
-
58
- | band | share of traffic | accuracy |
59
- |---|---|---|
60
- | ≥ 0.80 — act automatically | 62% | 93% |
61
- | 0.50–0.80 — confirm or review | 36% | 60% |
62
- | < 0.50 — route to a human | 2% | 32% |
63
-
64
- ## Confidence
65
-
66
- `confidence` is the model's probability that the answer it returned is the correct one:
67
-
68
- - **choice** — the probability of the selected option (`p_max`)
69
- - **score** — the probability mass that rounds to the reported expected level
70
- - **noul** — no confidence field; the probability *is* the answer
71
-
72
- This was chosen by measurement, not taste. Over 1,430 gold-labelled held-out rows,
73
- `p_max` beat normalized entropy, margin, chance-corrected top and collision entropy on
74
- **both** AUROC (0.816, ranking right answers above wrong ones) and ECE (0.036). Normalized
75
- entropy — the obvious first guess — was the worst of the lot: it is dominated by small
76
- probabilities, so it reads a decisive 0.85/0.08/0.07 as low confidence and routes 45% of
77
- traffic to human review at 64% accuracy.
78
-
79
  ## Usage
80
 
81
  The adapter is only meaningful with the exact prompt format it was trained on, so
@@ -155,6 +111,51 @@ a thinking budget the teacher commits to its answer with probability 1.0 on esse
155
  every row, which is useless for distillation — the measured mean entropy was 0.000 with
156
  reasoning versus 0.211 without.
157
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
158
  ## Limitations
159
 
160
  - **English only.**
 
18
 
19
  A QLoRA adapter that turns Gemma 4 E4B into a **System One** decision model: it reads a
20
  state, reads a typed question about it, and returns a calibrated probability distribution
21
+ over the allowed answers from a single **forward pass**, with no generation.
22
 
23
+ Since the answer is read from the logits at the option letters, it **cannot** answer
24
  outside the options it was given. There is no parsing, no retry loop, and no
25
  "as an AI language model".
26
 
 
32
  | `score` | rate against ordered levels | expected level, legend, probabilities, confidence |
33
  | `noul` | is this true? | a single probability |
34
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
35
  ## Usage
36
 
37
  The adapter is only meaningful with the exact prompt format it was trained on, so
 
111
  every row, which is useless for distillation — the measured mean entropy was 0.000 with
112
  reasoning versus 0.211 without.
113
 
114
+
115
+ ## Results
116
+
117
+ Evaluated on 1,898 held-out rows — 31 Super-NaturalInstructions **tasks** the model never
118
+ saw, 78 synthetic **states** it never saw, and SST-2 held out in its entirety.
119
+
120
+ | | accuracy | ECE | NLL | Brier |
121
+ |---|---|---|---|---|
122
+ | **overall** | **0.816** | **0.019** | 0.430 | 0.222 |
123
+ | choice (n=1196) | 0.809 | — | 0.431 | 0.238 |
124
+ | noul (n=547) | 0.819 | — | — | — |
125
+ | score (n=155) | 0.852 | — | — | — |
126
+
127
+ Score answers are off by **0.216 levels** on average (mean absolute error of the expected
128
+ level).
129
+
130
+ **Calibration is the point.** ECE of 0.019 means that when it reports 80% confidence it is
131
+ right about 80% of the time. Accuracy was flat from step 250 to the end of training while
132
+ ECE fell 0.086 → 0.019: the model did not learn to be right more often, it learned to be
133
+ honest about when it isn't.
134
+
135
+ Operationally, gating on confidence:
136
+
137
+ | band | share of traffic | accuracy |
138
+ |---|---|---|
139
+ | ≥ 0.80 — act automatically | 62% | 93% |
140
+ | 0.50–0.80 — confirm or review | 36% | 60% |
141
+ | < 0.50 — route to a human | 2% | 32% |
142
+
143
+ ## Confidence
144
+
145
+ `confidence` is the model's probability that the answer it returned is the correct one:
146
+
147
+ - **choice** — the probability of the selected option (`p_max`)
148
+ - **score** — the probability mass that rounds to the reported expected level
149
+ - **noul** — no confidence field; the probability *is* the answer
150
+
151
+ This was chosen by measurement, not taste. Over 1,430 gold-labelled held-out rows,
152
+ `p_max` beat normalized entropy, margin, chance-corrected top and collision entropy on
153
+ **both** AUROC (0.816, ranking right answers above wrong ones) and ECE (0.036). Normalized
154
+ entropy — the obvious first guess — was the worst of the lot: it is dominated by small
155
+ probabilities, so it reads a decisive 0.85/0.08/0.07 as low confidence and routes 45% of
156
+ traffic to human review at 64% accuracy.
157
+
158
+
159
  ## Limitations
160
 
161
  - **English only.**