jbrashear commited on
Commit
9aec78d
·
verified ·
1 Parent(s): df20afa

Temperatures: apply the hard fit to score questions (0.8329), keep train for choice and noul

Browse files

The train fit's score temperature (1.2162) was fitted to the smoothed ordinal target and made score probabilities too soft. Both calibration-split fits stay in temperatures.json under their names; the applied set is now choice 1.1863, noul 1.0903, score 0.8329, the rule the 27B has applied since 2026-09-26. Re-tempering the stored eval logits: question-weighted ECE over the 20 sets 0.080 to 0.071, NLL 0.589 to 0.580, no pick changes; both HelpSteer2 sets get worse. Card and eval/RESULTS.md ECE and Decision Score columns updated to match.

Files changed (4) hide show
  1. README.md +6 -6
  2. eval/RESULTS.md +35 -11
  3. jebadiah.json +1 -1
  4. temperatures.json +38 -2
README.md CHANGED
@@ -49,7 +49,7 @@ model-index:
49
  type: jevals-helpsteer2
50
  metrics:
51
  - type: decision_score_jevals
52
- value: 10.4
53
  - task:
54
  type: text-classification
55
  name: typed decisions (choice, noul, score)
@@ -69,7 +69,7 @@ Jebadiah (Jeb for short) is Frontier Infra's open System One style decision mode
69
 
70
  <img src="jeb-banner.png" alt="They call me Jeb. He does not talk much. He just decides." width="100%">
71
 
72
- This repository holds full bf16 weights: the v2 LoRA merged into `Qwen/Qwen3.5-9B` (revision `c2022362`, the chat checkpoint, thinking off). What changed from v1 is one thing, the base: v1 was trained on `Qwen/Qwen3.5-9B-Base`. The data, LoRA, objective and temperature fit are v1's.
73
 
74
  Made in Texas.
75
 
@@ -99,9 +99,9 @@ Measured by us with AINode's bench, one logit read per question, the same render
99
 
100
  The 27B is the same recipe on `Qwen/Qwen3.8-27B`; see [its card](https://huggingface.co/frontier-infra/jebadiah-27b). For scale, Bespoke's published table puts Nimble-9B at 74.8 and Jev at 76.0 on the same 13 Nimble public subsets, with their scorer.
101
 
102
- Where v2 is not better than v1: the Jevals HelpSteer2 Decision Score is 10.4 against 11.9, and its top level flips over identical repeats on 2.3% of questions against 0.3%. The Nimble public macro is flat (77.0): MultiNLI and PAWS give back what PubMedQA and HelpSteer2 gain. ECE on the Nimble 324 set rises from 0.063 to 0.085. The headline gain is 0.6 points, most of it from the Nimble 324 set. HelpSteer2 and SummEval are not zero-shot for Jeb: the pool trains on their train split and unscored articles (no evaluation item overlaps), so those rows are held-out items of a seen rubric.
103
 
104
- Every calibrated number applies the per-type temperatures in `temperatures.json`. Decision Scores, ECE, repeat flips, the in-distribution sets, every Nimble subset and the merge check are in [`eval/RESULTS.md`](eval/RESULTS.md); the per-question records are beside it in `eval/`. The nonce robustness pass is pending for this model, so no robustness number is claimed.
105
 
106
  ## Run it anywhere
107
 
@@ -169,7 +169,7 @@ These are independent builds. The agreement numbers under Other formats are for
169
 
170
  ## How it decides
171
 
172
- The prompt is AINode's own decide rendering (source commit `e5c08938`, hash in `prompt_contract.json`), through the chat template with thinking off. The option labels are single tokens; the answer is the distribution over those label tokens at the last prompt position, read in fp32 and temperature scaled per type (`temperatures.json`: choice 1.19, noul 1.09, score 1.22). Nothing is generated. Those are the `train` fit, which minimises NLL against the soft or ordinal target the model was taught; it softens, and it costs some ECE against hard labels on the calibration split. A sharper `hard` fit (0.79 / 0.67 / 0.83) is in the same file for a consumer who gates on the argmax. Thresholds belong to the caller: act on a high probability, confirm or escalate on a middle one, hand a low one to a person or a bigger model. The model never refuses.
173
 
174
  ## Training
175
 
@@ -184,7 +184,7 @@ The prompt is AINode's own decide rendering (source commit `e5c08938`, hash in `
184
  - Single-hop judgments only. Split a chain of inference into hops.
185
  - A choice question is capped at 20 options on `/v1/systemone`; Banking77's 77 options were scored with an extended single-token alphabet for the benchmark only.
186
  - English data. Training cut states past 2,048 prompt tokens.
187
- - Calibration was fitted on the training distribution. Refit before trusting a threshold.
188
 
189
  ## Versioning and license
190
 
 
49
  type: jevals-helpsteer2
50
  metrics:
51
  - type: decision_score_jevals
52
+ value: 10.5
53
  - task:
54
  type: text-classification
55
  name: typed decisions (choice, noul, score)
 
69
 
70
  <img src="jeb-banner.png" alt="They call me Jeb. He does not talk much. He just decides." width="100%">
71
 
72
+ This repository holds full bf16 weights: the v2 LoRA merged into `Qwen/Qwen3.5-9B` (revision `c2022362`, the chat checkpoint, thinking off). What changed from v1 is one thing, the base: v1 was trained on `Qwen/Qwen3.5-9B-Base`. The data, LoRA, objective and temperature fits are v1's; since 2026-09-29 the score temperature applied is the fit against the label (see How it decides).
73
 
74
  Made in Texas.
75
 
 
99
 
100
  The 27B is the same recipe on `Qwen/Qwen3.8-27B`; see [its card](https://huggingface.co/frontier-infra/jebadiah-27b). For scale, Bespoke's published table puts Nimble-9B at 74.8 and Jev at 76.0 on the same 13 Nimble public subsets, with their scorer.
101
 
102
+ Where v2 is not better than v1: the Jevals HelpSteer2 Decision Score is 10.5 against 11.9, and its top level flips over identical repeats on 2.3% of questions against 0.3%. The Nimble public macro is flat (77.0): MultiNLI and PAWS give back what PubMedQA and HelpSteer2 gain. ECE on the Nimble 324 set rises from 0.063 to 0.070. The headline gain is 0.6 points, most of it from the Nimble 324 set. HelpSteer2 and SummEval are not zero-shot for Jeb: the pool trains on their train split and unscored articles (no evaluation item overlaps), so those rows are held-out items of a seen rubric.
103
 
104
+ Every calibrated number applies the per-type temperatures in `temperatures.json`: choice 1.19 and noul 1.09 from the fit to the training target, score 0.83 from the fit to the label, all three on the held-out calibration split. The score temperature was refit on 2026-09-29 from 1.22 (fitted to the smoothed ordinal training target) to 0.83, the rule the 27B has used since 2026-09-26, with no change to any pick. It lowers ECE on typed-decisions from 0.156 to 0.121 and on Nimble 324 from 0.085 to 0.070, and raises it on Jevals HelpSteer2 from 0.039 to 0.082 and Nimble public HelpSteer2 from 0.073 to 0.098; the comparison is in [`eval/RESULTS.md`](eval/RESULTS.md#temperature-fits). Decision Scores, ECE, repeat flips, the in-distribution sets, every Nimble subset and the merge check are in [`eval/RESULTS.md`](eval/RESULTS.md); the per-question records are beside it in `eval/`. The nonce robustness pass is pending for this model, so no robustness number is claimed.
105
 
106
  ## Run it anywhere
107
 
 
169
 
170
  ## How it decides
171
 
172
+ The prompt is AINode's own decide rendering (source commit `e5c08938`, hash in `prompt_contract.json`), through the chat template with thinking off. The option labels are single tokens; the answer is the distribution over those label tokens at the last prompt position, read in fp32 and temperature scaled per type (`temperatures.json`: choice 1.19, noul 1.09, score 0.83). Nothing is generated. Choice and noul use the `train` fit (NLL against the soft target the model was taught; it softens), score uses the `hard` fit (NLL against the label; it sharpens), because the ordinal score target is smoothed on purpose and a temperature fitted to it made score probabilities too soft. Both complete fits, `train` (1.19 / 1.09 / 1.22) and `hard` (0.79 / 0.67 / 0.83), are in the same file. Thresholds belong to the caller: act on a high probability, confirm or escalate on a middle one, hand a low one to a person or a bigger model. The model never refuses.
173
 
174
  ## Training
175
 
 
184
  - Single-hop judgments only. Split a chain of inference into hops.
185
  - A choice question is capped at 20 options on `/v1/systemone`; Banking77's 77 options were scored with an extended single-token alphabet for the benchmark only.
186
  - English data. Training cut states past 2,048 prompt tokens.
187
+ - Calibration was fitted on the training distribution's calibration split, where HelpSteer2's labels are close to uniform. On traffic skewed like natural HelpSteer2 (over 70% of labels at 3 or 4) score answers are too confident: ECE on Jevals HelpSteer2 is 0.082. Refit on your own labelled traffic before trusting a threshold ([how](eval/RESULTS.md#temperature-fits)).
188
 
189
  ## Versioning and license
190
 
eval/RESULTS.md CHANGED
@@ -16,11 +16,11 @@ The v1 columns are the published [v1](https://huggingface.co/frontier-infra/jeba
16
  |---|---|---:|---:|---:|---:|---:|---:|---:|
17
  | Jevals PubMedQA (300) | noul | 90.3 | 89.7 | 62.0 | 68.6 | 66.8 | 0.025 | 0.0% |
18
  | Jevals Banking77 (300, 77 options) | choice | 70.7 | 70.0 | 1.3 | 58.2 | 56.9 | 0.076 | 0.7% |
19
- | Jevals HelpSteer2 helpfulness (300, 5 levels) | score | 40.7 | 40.3 | 41.7 | 10.4 | 11.9 | 0.039 (raw 0.066) | 2.3% |
20
- | Nimble held-out eval (324) | mixed | 81.2 | 78.7 | 17.6 | 65.0 | 65.0 | 0.085 | 0.3% |
21
- | Kev transfer-v4 test (764) | mixed | 83.8 | 84.0 | 21.5 | 67.9 | 67.0 | 0.025 | 0.3% |
22
- | Kev decision-v7 test (1,440) | mixed | 81.2 | 80.8 | 20.3 | 71.4 | 70.6 | 0.050 | n/a |
23
- | typed-decisions test (2,000) | mixed | 78.3 | 78.9 | 15.3 | 56.5 | 56.4 | 0.156 | 0.1% |
24
 
25
  Nimble's 13 public human-labelled subsets (3,880 questions), macro accuracy: **77.0** (v1: 77.0).
26
  For scale, Bespoke's published table puts Nimble-9B at 74.8 and Jev at 76.0 on the same subsets with their scorer.
@@ -30,18 +30,18 @@ For scale, Bespoke's published table puts Nimble-9B at 74.8 and Jev at 76.0 on t
30
  | aegis2 (250) | 78.4 | 79.2 | 38.0 | 35.8 |
31
  | boolq (300) | 86.0 | 85.7 | 54.6 | 57.7 |
32
  | civil_comments (300) | 86.0 | 86.7 | -7.4 | -4.8 |
33
- | helpsteer2 (249) | 43.4 | 41.0 | 13.3 | 15.2 |
34
  | massive-de-DE (350) | 84.6 | 84.9 | 74.3 | 73.3 |
35
  | massive-en-US (350) | 85.4 | 85.7 | 77.3 | 76.0 |
36
  | multinli (299) | 84.9 | 87.6 | 67.2 | 70.6 |
37
  | paws (250) | 83.6 | 85.6 | 54.0 | 57.2 |
38
  | pubmedqa (250) | 72.8 | 68.8 | 38.5 | 34.5 |
39
  | squad2 (299) | 79.9 | 80.3 | 42.9 | 43.6 |
40
- | summeval-consistency (144) | 87.5 | 87.5 | 42.6 | 36.0 |
41
- | summeval-relevance (240) | 52.5 | 51.7 | 23.4 | 25.9 |
42
  | vitaminc-dev (599) | 75.6 | 76.3 | 41.8 | 39.0 |
43
 
44
- Where v2 is not better. Jevals HelpSteer2 Decision Score is 10.4 against v1's 11.9, and its top level flips over identical repeats on 2.3% of questions against 0.3%. The Nimble public macro is flat (77.0 against 77.0): MultiNLI (84.9 against 87.6) and PAWS (83.6 against 85.6) give back what PubMedQA (72.8 against 68.8) and HelpSteer2 gain. ECE on the Nimble 324 set rises from 0.063 to 0.085. The headline gain is 0.6 points, most of it from the Nimble 324 set (81.2 against 78.7).
45
 
46
  Two things to read carefully, unchanged from v1 because the pool is unchanged. HelpSteer2 and SummEval are not
47
  zero-shot for Jeb: the pool trains on the HelpSteer2 **train** split (the Jevals and Nimble items come from
@@ -50,7 +50,10 @@ are held out entirely), so those rows are held-out items of a seen rubric. And t
50
  (typed-decisions, Kev decision-v7) are the test splits of sources in the pool.
51
 
52
  Every eval record (per question: option keys, probabilities, pick, label, repeat, option order) is in
53
- `eval/`, with `eval/results.json` carrying the full metric set. The records were scored on the unmerged adapter;
 
 
 
54
  the merged weights here reproduce them (see Merge verification below).
55
 
56
  ## Robustness to irrelevant content (nonce test)
@@ -59,7 +62,28 @@ The nonce pass (a fresh UUID planted in the state or appended to the instruction
59
 
60
  ## Temperature fits
61
 
62
- `temperatures.json` carries two fits on the calibration split (689 records, 921 questions: choice 225, noul 147, score 549). The `hard` fit minimises NLL against the argmax label and sharpens (T 0.79 / 0.67 / 0.83). The `train` fit minimises NLL against the training target, the soft or ordinal distribution the model was taught, and softens (T 1.19 / 1.09 / 1.22). v2 applies `train`, as v1 does, because the ordinal target is what makes rubric probabilities honest, and the `hard` fit is exactly the sharpening that made v0's HelpSteer2 calibration worse. The price is visible in the same file: on the calibration split measured against hard labels, ECE moves from 0.074 to 0.089 (choice), 0.064 to 0.064 (noul) and 0.092 to 0.128 (score) while NLL on the training target falls (0.4594 to 0.4542, 0.3776 to 0.3768, 1.1017 to 1.0929). The ECE column in the results table is measured on the evaluation sets, not on this split, with the applied temperatures. Both fits are in the file; a consumer who gates on the argmax may prefer `hard`, and either way should refit on its own data.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
63
 
64
  ## Merge verification
65
 
 
16
  |---|---|---:|---:|---:|---:|---:|---:|---:|
17
  | Jevals PubMedQA (300) | noul | 90.3 | 89.7 | 62.0 | 68.6 | 66.8 | 0.025 | 0.0% |
18
  | Jevals Banking77 (300, 77 options) | choice | 70.7 | 70.0 | 1.3 | 58.2 | 56.9 | 0.076 | 0.7% |
19
+ | Jevals HelpSteer2 helpfulness (300, 5 levels) | score | 40.7 | 40.3 | 41.7 | 10.5 | 11.9 | 0.082 (raw 0.066) | 2.3% |
20
+ | Nimble held-out eval (324) | mixed | 81.2 | 78.7 | 17.6 | 65.0 | 65.0 | 0.070 (raw 0.059) | 0.3% |
21
+ | Kev transfer-v4 test (764) | mixed | 83.8 | 84.0 | 21.5 | 67.9 | 67.0 | 0.027 | 0.3% |
22
+ | Kev decision-v7 test (1,440) | mixed | 81.2 | 80.8 | 20.3 | 71.4 | 70.6 | 0.048 | n/a |
23
+ | typed-decisions test (2,000) | mixed | 78.3 | 78.9 | 15.3 | 57.1 | 56.4 | 0.121 | 0.1% |
24
 
25
  Nimble's 13 public human-labelled subsets (3,880 questions), macro accuracy: **77.0** (v1: 77.0).
26
  For scale, Bespoke's published table puts Nimble-9B at 74.8 and Jev at 76.0 on the same subsets with their scorer.
 
30
  | aegis2 (250) | 78.4 | 79.2 | 38.0 | 35.8 |
31
  | boolq (300) | 86.0 | 85.7 | 54.6 | 57.7 |
32
  | civil_comments (300) | 86.0 | 86.7 | -7.4 | -4.8 |
33
+ | helpsteer2 (249) | 43.4 | 41.0 | 12.3 | 15.2 |
34
  | massive-de-DE (350) | 84.6 | 84.9 | 74.3 | 73.3 |
35
  | massive-en-US (350) | 85.4 | 85.7 | 77.3 | 76.0 |
36
  | multinli (299) | 84.9 | 87.6 | 67.2 | 70.6 |
37
  | paws (250) | 83.6 | 85.6 | 54.0 | 57.2 |
38
  | pubmedqa (250) | 72.8 | 68.8 | 38.5 | 34.5 |
39
  | squad2 (299) | 79.9 | 80.3 | 42.9 | 43.6 |
40
+ | summeval-consistency (144) | 87.5 | 87.5 | 49.4 | 36.0 |
41
+ | summeval-relevance (240) | 52.5 | 51.7 | 23.9 | 25.9 |
42
  | vitaminc-dev (599) | 75.6 | 76.3 | 41.8 | 39.0 |
43
 
44
+ Where v2 is not better. Jevals HelpSteer2 Decision Score is 10.5 against v1's 11.9, and its top level flips over identical repeats on 2.3% of questions against 0.3%. The Nimble public macro is flat (77.0 against 77.0): MultiNLI (84.9 against 87.6) and PAWS (83.6 against 85.6) give back what PubMedQA (72.8 against 68.8) and HelpSteer2 gain. ECE on the Nimble 324 set rises from 0.063 to 0.070. The headline gain is 0.6 points, most of it from the Nimble 324 set (81.2 against 78.7).
45
 
46
  Two things to read carefully, unchanged from v1 because the pool is unchanged. HelpSteer2 and SummEval are not
47
  zero-shot for Jeb: the pool trains on the HelpSteer2 **train** split (the Jevals and Nimble items come from
 
50
  (typed-decisions, Kev decision-v7) are the test splits of sources in the pool.
51
 
52
  Every eval record (per question: option keys, probabilities, pick, label, repeat, option order) is in
53
+ `eval/`, with `eval/results.json` carrying the full metric set. The records were scored on the unmerged adapter with the
54
+ previous temperatures (each file's `model.temperatures`); the ECE and Decision Score columns above re-temper those
55
+ same logits to the temperatures now in `temperatures.json` (see Temperature fits), so they differ from `results.json`
56
+ on the sets with score questions. Accuracy is unchanged;
57
  the merged weights here reproduce them (see Merge verification below).
58
 
59
  ## Robustness to irrelevant content (nonce test)
 
62
 
63
  ## Temperature fits
64
 
65
+ `temperatures.json` carries two fits on the calibration split (689 records, 921 questions: choice 225, noul 147, score 549), both kept under their names. The `hard` fit minimises NLL against the argmax label and sharpens (T 0.79 / 0.67 / 0.83). The `train` fit minimises NLL against the training target, the soft or ordinal distribution the model was taught, and softens (T 1.19 / 1.09 / 1.22).
66
+
67
+ **Applied since 2026-09-29: `mixed`, choice 1.19 and noul 1.09 from `train`, score 0.83 from `hard`**, the rule the [27B](https://huggingface.co/frontier-infra/jebadiah-27b/blob/main/eval/RESULTS.md#temperature-fits) has applied since 2026-09-26. 9B v2 first shipped with `train` for all three types. The `train` score target is the ordinal kernel, a deliberately smoothed distribution, so fitting T to it softens past what the labels support. No model was retrained and nothing was fitted on an evaluation set: both numbers are the calibration split's own fits. On the calibration split itself, fitting on one half and scoring the other (200 random splits by record family; raw logits for its 549 score questions read through AINode on our fleet, which reproduce the recorded `hard` score fit, 0.837 against 0.833), score ECE is 0.119 with `train`, 0.093 with no temperature and 0.079 with `hard`; NLL 0.953, 0.932, 0.934.
68
+
69
+ Over all 20 evaluation sets (re-tempering the stored per-question logits), question-weighted ECE 0.080 with `train`, 0.079 with no temperature, 0.089 with `hard`, 0.071 applied; macro ECE 0.071, 0.074, 0.093, 0.065; NLL 0.589, 0.591, 0.634, 0.580; macro Decision Score (Jevals) 47.9, 47.8, 46.1, 48.3. Only the sets with score questions change:
70
+
71
+ | Set (score questions present) | ECE, none | ECE, `train` (previous) | ECE, `hard` | ECE, applied | NLL, `train` | NLL, applied |
72
+ |---|---:|---:|---:|---:|---:|---:|
73
+ | Jevals HelpSteer2 (score) | 0.066 | 0.039 | 0.082 | 0.082 | 1.308 | 1.314 |
74
+ | Nimble 324 (mixed) | 0.059 | 0.085 | 0.037 | 0.070 | 0.490 | 0.481 |
75
+ | Kev transfer-v4 (mixed) | 0.037 | 0.025 | 0.067 | 0.027 | 0.442 | 0.440 |
76
+ | Kev decision-v7 (mixed) | 0.059 | 0.050 | 0.087 | 0.048 | 0.530 | 0.534 |
77
+ | typed-decisions (mixed) | 0.125 | 0.156 | 0.078 | 0.121 | 0.591 | 0.553 |
78
+ | Nimble public HelpSteer2 (score) | 0.072 | 0.073 | 0.098 | 0.098 | 1.272 | 1.292 |
79
+ | Nimble public SummEval consistency (score) | 0.149 | 0.195 | 0.118 | 0.118 | 0.532 | 0.462 |
80
+ | Nimble public SummEval relevance (score) | 0.044 | 0.084 | 0.033 | 0.033 | 1.098 | 1.080 |
81
+
82
+ It is not a win everywhere: both HelpSteer2 sets were better calibrated with `train`. We checked whether a separate score temperature per rubric, fitted on the calibration split, would keep HelpSteer2 where it was. It does not, because the split's own HelpSteer2 questions also fit about 0.85: their labels are close to uniform across 0 to 4, while the Jevals and Nimble HelpSteer2 items put over 70% of their labels at 3 or 4, where the model is less accurate and too confident. That is a shift in the label mix, which no fit on this split can see. Per-rubric fits also overfit a split this small (its SummEval part is four articles, and a 17-question rubric's fit ranged from 0.8 to 15 across halves), so one temperature per type stays.
83
+
84
+ Calibration was fitted on the training distribution; refit the temperatures on your own labelled traffic before trusting a threshold. Log the raw distribution (`"calibration": "raw"` on AINode, `--no-temperatures` in the scripts), label a random sample of a few hundred decisions, fit one T per question type against the label on half of them and check ECE and NLL on the other half before you replace the file.
85
+
86
+ ECE here uses 10 equal-width bins on the pick's probability rounded to a whole percent (bin = round(100p) // 10), repeat 0. Binning at floor(10p) gives slightly different values on the flat score sets (on the 9B's Nimble public HelpSteer2 records with `train`, 0.060 against 0.073).
87
 
88
  ## Merge verification
89
 
jebadiah.json CHANGED
@@ -74,7 +74,7 @@
74
  "temperatures": {
75
  "choice": 1.1863,
76
  "noul": 1.0903,
77
- "score": 1.2162
78
  },
79
  "note": "same layout, config and tokenizer as the base repo; only the LoRA-targeted language_model weights differ. The served route applies no temperature."
80
  }
 
74
  "temperatures": {
75
  "choice": 1.1863,
76
  "noul": 1.0903,
77
+ "score": 0.8329
78
  },
79
  "note": "same layout, config and tokenizer as the base repo; only the LoRA-targeted language_model weights differ. The served route applies no temperature."
80
  }
temperatures.json CHANGED
@@ -2,9 +2,25 @@
2
  "temperatures": {
3
  "choice": 1.1863,
4
  "noul": 1.0903,
5
- "score": 1.2162
6
  },
7
- "applied_target": "train",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8
  "calib_file": "/workspace/jeb/data-v1/calib.jsonl",
9
  "n": {
10
  "choice": 225,
@@ -45,6 +61,26 @@
45
  "nll_before": 1.1017,
46
  "nll_after": 1.0929
47
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
48
  }
49
  },
50
  "nll_before": {
 
2
  "temperatures": {
3
  "choice": 1.1863,
4
  "noul": 1.0903,
5
+ "score": 0.8329
6
  },
7
+ "applied_target": "mixed",
8
+ "applied_fits": {
9
+ "choice": "train",
10
+ "noul": "train",
11
+ "score": "hard"
12
+ },
13
+ "previous": {
14
+ "applied_target": "train",
15
+ "temperatures": {
16
+ "choice": 1.1863,
17
+ "noul": 1.0903,
18
+ "score": 1.2162
19
+ },
20
+ "replaced": "2026-09-29"
21
+ },
22
+ "why": "Refit 2026-09-29 without retraining, the rule the 27B has applied since 2026-09-26. Both fits are unchanged and come from the calibration split only. Score questions calibrate better at the hard fit (T 0.83) than at the train fit (T 1.22), whose ordinal target is deliberately smoothed: on held-out halves of the calibration split score ECE 0.119 -> 0.079. Choice and noul stay on the train fit. Re-tempering the stored logits of the 20 evaluation sets: question-weighted ECE 0.0796 -> 0.0709, macro 0.0709 -> 0.0653, NLL 0.5887 -> 0.5799; accuracy unchanged. Jevals HelpSteer2 gets worse (0.039 -> 0.082). See eval/RESULTS.md, Temperature fits.",
23
+ "top_level_stats_describe": "the train fit (nll_before/nll_after/ece_before/ece_after below are its calibration-split numbers)",
24
  "calib_file": "/workspace/jeb/data-v1/calib.jsonl",
25
  "n": {
26
  "choice": 225,
 
61
  "nll_before": 1.1017,
62
  "nll_after": 1.0929
63
  }
64
+ },
65
+ "mixed": {
66
+ "choice": {
67
+ "T": 1.1863,
68
+ "nll_before": 0.4594,
69
+ "nll_after": 0.4542,
70
+ "source": "train"
71
+ },
72
+ "noul": {
73
+ "T": 1.0903,
74
+ "nll_before": 0.3776,
75
+ "nll_after": 0.3768,
76
+ "source": "train"
77
+ },
78
+ "score": {
79
+ "T": 0.8329,
80
+ "nll_before": 0.9298,
81
+ "nll_after": 0.9225,
82
+ "source": "hard"
83
+ }
84
  }
85
  },
86
  "nll_before": {