colesmcintosh commited on
Commit
bd8e994
·
verified ·
1 Parent(s): d2a855e

Add measured confidence

Browse files
Files changed (1) hide show
  1. README.md +40 -4
README.md CHANGED
@@ -126,9 +126,9 @@ APIs like [TypeSafe's](https://docs.typesafe.ai/api) ask typed `choice`, `noul`,
126
 
127
  | Question type | mev options | Answer |
128
  |---|---|---|
129
- | `choice` with `criteria: {key: description}` | One option per criteria entry | Most likely option, plus a probability per option |
130
  | `noul` (yes/no) | `A` = yes, `B` = no | Probability of `A` |
131
- | `score` with ordered `criteria: [...]` | One option per level, in order | Probability-weighted mean of the level indexes |
132
 
133
  ```python
134
  import math
@@ -153,13 +153,49 @@ def option_probabilities(model, tokenizer, prompt, options):
153
  scores = {o["key"]: logits[tokenizer.convert_tokens_to_ids(o["label"])].item() for o in options}
154
  total = sum(math.exp(s) for s in scores.values())
155
  return {key: math.exp(s) / total for key, s in scores.items()}
 
 
 
156
  ```
157
 
158
  Compared with a typed-question API:
159
 
160
  - **One question per call.** mev was trained on single decisions. To ask several questions about the same state, make one call per question, in parallel if you like.
161
  - **2 to 24 options.** That is the range seen in training. Larger option sets are untested.
162
- - **No `confidence` field, and probabilities are not calibrated.** Letter probabilities show which option the model prefers, but 0.9 does not mean the answer is right 90% of the time. If you derive a confidence value, check it against your own labeled data first.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
163
 
164
  ## Training
165
 
@@ -213,7 +249,7 @@ These are development sets that shaped the recipe, not an independent benchmark.
213
 
214
  - Generic chat is not the intended interface and may produce prose.
215
  - The model can be wrong. Do not use it as the only authority for high-impact decisions.
216
- - Prompt injection, multilingual behavior, calibration, and out-of-distribution robustness have not been comprehensively evaluated.
217
 
218
  ## License
219
 
 
126
 
127
  | Question type | mev options | Answer |
128
  |---|---|---|
129
+ | `choice` with `criteria: {key: description}` | One option per criteria entry | Most likely option, a probability per option, and `confidence` |
130
  | `noul` (yes/no) | `A` = yes, `B` = no | Probability of `A` |
131
+ | `score` with ordered `criteria: [...]` | One option per level, in order | Probability-weighted mean of the level indexes, and `confidence` |
132
 
133
  ```python
134
  import math
 
153
  scores = {o["key"]: logits[tokenizer.convert_tokens_to_ids(o["label"])].item() for o in options}
154
  total = sum(math.exp(s) for s in scores.values())
155
  return {key: math.exp(s) / total for key, s in scores.items()}
156
+
157
+ def confidence(probabilities):
158
+ return max(probabilities.values())
159
  ```
160
 
161
  Compared with a typed-question API:
162
 
163
  - **One question per call.** mev was trained on single decisions. To ask several questions about the same state, make one call per question, in parallel if you like.
164
  - **2 to 24 options.** That is the range seen in training. Larger option sets are untested.
165
+ - **`confidence` is computed client-side.** The model doesn't return a confidence field. Take the probability of the chosen option; see [Confidence](#confidence) for how well it matches real accuracy.
166
+
167
+ ## Confidence
168
+
169
+ mev's `confidence` is the probability of the chosen option, after normalizing over the task's option letters. On the main development set (1,000 decisions), it closely tracks real accuracy, with an expected calibration error of 0.038. It also separates right answers from wrong ones well, with an AUROC of 0.88.
170
+
171
+ Accuracy when acting only on answers above a confidence threshold:
172
+
173
+ | Confidence at or above | Share of decisions kept | Accuracy on those |
174
+ |---:|---:|---:|
175
+ | 0.50 | 92.7% | 89.0% |
176
+ | 0.70 | 77.1% | 96.2% |
177
+ | 0.80 | 69.3% | 97.4% |
178
+ | 0.90 | 60.5% | 98.0% |
179
+ | 0.95 | 50.6% | 98.6% |
180
+ | 0.99 | 31.3% | 100.0% |
181
+
182
+ Reliability by confidence band:
183
+
184
+ | Confidence band | Decisions | Mean confidence | Accuracy |
185
+ |---|---:|---:|---:|
186
+ | below 0.50 | 73 | 0.44 | 57.5% |
187
+ | 0.50 to 0.70 | 156 | 0.60 | 53.2% |
188
+ | 0.70 to 0.80 | 78 | 0.75 | 85.9% |
189
+ | 0.80 to 0.90 | 88 | 0.85 | 93.2% |
190
+ | 0.90 to 0.95 | 99 | 0.93 | 94.9% |
191
+ | 0.95 to 0.99 | 193 | 0.97 | 96.4% |
192
+ | 0.99 and above | 313 | 1.00 | 100.0% |
193
+
194
+ Between 0.7 and 0.95, mev is slightly underconfident: it is right more often than its confidence says. Between 0.5 and 0.7, it is overconfident. A common pattern is to act automatically above a threshold like 0.9 and send the rest to a fallback, such as a larger model or a person.
195
+
196
+ Other ways to compute confidence were measured too. The gap between the top two options ranks answers about as well (AUROC 0.88) but has a calibration error of 0.14. A score based on entropy, which measures how spread out the probabilities are, does worse on both (AUROC 0.85, calibration error 0.17).
197
+
198
+ These numbers come from development sets that shaped the training recipe. Calibration can shift on your own data, so check your threshold against a labeled sample before relying on it.
199
 
200
  ## Training
201
 
 
249
 
250
  - Generic chat is not the intended interface and may produce prose.
251
  - The model can be wrong. Do not use it as the only authority for high-impact decisions.
252
+ - Prompt injection, multilingual behavior, and out-of-distribution robustness have not been comprehensively evaluated. Calibration was measured only on the development sets.
253
 
254
  ## License
255