Instructions to use frontier-infra/jebadiah-9b-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use frontier-infra/jebadiah-9b-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="frontier-infra/jebadiah-9b-v2") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("frontier-infra/jebadiah-9b-v2") model = AutoModelForMultimodalLM.from_pretrained("frontier-infra/jebadiah-9b-v2", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use frontier-infra/jebadiah-9b-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "frontier-infra/jebadiah-9b-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "frontier-infra/jebadiah-9b-v2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/frontier-infra/jebadiah-9b-v2
- SGLang
How to use frontier-infra/jebadiah-9b-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "frontier-infra/jebadiah-9b-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "frontier-infra/jebadiah-9b-v2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "frontier-infra/jebadiah-9b-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "frontier-infra/jebadiah-9b-v2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use frontier-infra/jebadiah-9b-v2 with Docker Model Runner:
docker model run hf.co/frontier-infra/jebadiah-9b-v2
Temperatures: apply the hard fit to score questions (0.8329), keep train for choice and noul
Browse filesThe train fit's score temperature (1.2162) was fitted to the smoothed ordinal target and made score probabilities too soft. Both calibration-split fits stay in temperatures.json under their names; the applied set is now choice 1.1863, noul 1.0903, score 0.8329, the rule the 27B has applied since 2026-09-26. Re-tempering the stored eval logits: question-weighted ECE over the 20 sets 0.080 to 0.071, NLL 0.589 to 0.580, no pick changes; both HelpSteer2 sets get worse. Card and eval/RESULTS.md ECE and Decision Score columns updated to match.
- README.md +6 -6
- eval/RESULTS.md +35 -11
- jebadiah.json +1 -1
- temperatures.json +38 -2
|
@@ -49,7 +49,7 @@ model-index:
|
|
| 49 |
type: jevals-helpsteer2
|
| 50 |
metrics:
|
| 51 |
- type: decision_score_jevals
|
| 52 |
-
value: 10.
|
| 53 |
- task:
|
| 54 |
type: text-classification
|
| 55 |
name: typed decisions (choice, noul, score)
|
|
@@ -69,7 +69,7 @@ Jebadiah (Jeb for short) is Frontier Infra's open System One style decision mode
|
|
| 69 |
|
| 70 |
<img src="jeb-banner.png" alt="They call me Jeb. He does not talk much. He just decides." width="100%">
|
| 71 |
|
| 72 |
-
This repository holds full bf16 weights: the v2 LoRA merged into `Qwen/Qwen3.5-9B` (revision `c2022362`, the chat checkpoint, thinking off). What changed from v1 is one thing, the base: v1 was trained on `Qwen/Qwen3.5-9B-Base`. The data, LoRA, objective and temperature
|
| 73 |
|
| 74 |
Made in Texas.
|
| 75 |
|
|
@@ -99,9 +99,9 @@ Measured by us with AINode's bench, one logit read per question, the same render
|
|
| 99 |
|
| 100 |
The 27B is the same recipe on `Qwen/Qwen3.8-27B`; see [its card](https://huggingface.co/frontier-infra/jebadiah-27b). For scale, Bespoke's published table puts Nimble-9B at 74.8 and Jev at 76.0 on the same 13 Nimble public subsets, with their scorer.
|
| 101 |
|
| 102 |
-
Where v2 is not better than v1: the Jevals HelpSteer2 Decision Score is 10.
|
| 103 |
|
| 104 |
-
Every calibrated number applies the per-type temperatures in `temperatures.json`. Decision Scores, ECE, repeat flips, the in-distribution sets, every Nimble subset and the merge check are in [`eval/RESULTS.md`](eval/RESULTS.md); the per-question records are beside it in `eval/`. The nonce robustness pass is pending for this model, so no robustness number is claimed.
|
| 105 |
|
| 106 |
## Run it anywhere
|
| 107 |
|
|
@@ -169,7 +169,7 @@ These are independent builds. The agreement numbers under Other formats are for
|
|
| 169 |
|
| 170 |
## How it decides
|
| 171 |
|
| 172 |
-
The prompt is AINode's own decide rendering (source commit `e5c08938`, hash in `prompt_contract.json`), through the chat template with thinking off. The option labels are single tokens; the answer is the distribution over those label tokens at the last prompt position, read in fp32 and temperature scaled per type (`temperatures.json`: choice 1.19, noul 1.09, score
|
| 173 |
|
| 174 |
## Training
|
| 175 |
|
|
@@ -184,7 +184,7 @@ The prompt is AINode's own decide rendering (source commit `e5c08938`, hash in `
|
|
| 184 |
- Single-hop judgments only. Split a chain of inference into hops.
|
| 185 |
- A choice question is capped at 20 options on `/v1/systemone`; Banking77's 77 options were scored with an extended single-token alphabet for the benchmark only.
|
| 186 |
- English data. Training cut states past 2,048 prompt tokens.
|
| 187 |
-
- Calibration was fitted on the training distribution. Refit before trusting a threshold.
|
| 188 |
|
| 189 |
## Versioning and license
|
| 190 |
|
|
|
|
| 49 |
type: jevals-helpsteer2
|
| 50 |
metrics:
|
| 51 |
- type: decision_score_jevals
|
| 52 |
+
value: 10.5
|
| 53 |
- task:
|
| 54 |
type: text-classification
|
| 55 |
name: typed decisions (choice, noul, score)
|
|
|
|
| 69 |
|
| 70 |
<img src="jeb-banner.png" alt="They call me Jeb. He does not talk much. He just decides." width="100%">
|
| 71 |
|
| 72 |
+
This repository holds full bf16 weights: the v2 LoRA merged into `Qwen/Qwen3.5-9B` (revision `c2022362`, the chat checkpoint, thinking off). What changed from v1 is one thing, the base: v1 was trained on `Qwen/Qwen3.5-9B-Base`. The data, LoRA, objective and temperature fits are v1's; since 2026-09-29 the score temperature applied is the fit against the label (see How it decides).
|
| 73 |
|
| 74 |
Made in Texas.
|
| 75 |
|
|
|
|
| 99 |
|
| 100 |
The 27B is the same recipe on `Qwen/Qwen3.8-27B`; see [its card](https://huggingface.co/frontier-infra/jebadiah-27b). For scale, Bespoke's published table puts Nimble-9B at 74.8 and Jev at 76.0 on the same 13 Nimble public subsets, with their scorer.
|
| 101 |
|
| 102 |
+
Where v2 is not better than v1: the Jevals HelpSteer2 Decision Score is 10.5 against 11.9, and its top level flips over identical repeats on 2.3% of questions against 0.3%. The Nimble public macro is flat (77.0): MultiNLI and PAWS give back what PubMedQA and HelpSteer2 gain. ECE on the Nimble 324 set rises from 0.063 to 0.070. The headline gain is 0.6 points, most of it from the Nimble 324 set. HelpSteer2 and SummEval are not zero-shot for Jeb: the pool trains on their train split and unscored articles (no evaluation item overlaps), so those rows are held-out items of a seen rubric.
|
| 103 |
|
| 104 |
+
Every calibrated number applies the per-type temperatures in `temperatures.json`: choice 1.19 and noul 1.09 from the fit to the training target, score 0.83 from the fit to the label, all three on the held-out calibration split. The score temperature was refit on 2026-09-29 from 1.22 (fitted to the smoothed ordinal training target) to 0.83, the rule the 27B has used since 2026-09-26, with no change to any pick. It lowers ECE on typed-decisions from 0.156 to 0.121 and on Nimble 324 from 0.085 to 0.070, and raises it on Jevals HelpSteer2 from 0.039 to 0.082 and Nimble public HelpSteer2 from 0.073 to 0.098; the comparison is in [`eval/RESULTS.md`](eval/RESULTS.md#temperature-fits). Decision Scores, ECE, repeat flips, the in-distribution sets, every Nimble subset and the merge check are in [`eval/RESULTS.md`](eval/RESULTS.md); the per-question records are beside it in `eval/`. The nonce robustness pass is pending for this model, so no robustness number is claimed.
|
| 105 |
|
| 106 |
## Run it anywhere
|
| 107 |
|
|
|
|
| 169 |
|
| 170 |
## How it decides
|
| 171 |
|
| 172 |
+
The prompt is AINode's own decide rendering (source commit `e5c08938`, hash in `prompt_contract.json`), through the chat template with thinking off. The option labels are single tokens; the answer is the distribution over those label tokens at the last prompt position, read in fp32 and temperature scaled per type (`temperatures.json`: choice 1.19, noul 1.09, score 0.83). Nothing is generated. Choice and noul use the `train` fit (NLL against the soft target the model was taught; it softens), score uses the `hard` fit (NLL against the label; it sharpens), because the ordinal score target is smoothed on purpose and a temperature fitted to it made score probabilities too soft. Both complete fits, `train` (1.19 / 1.09 / 1.22) and `hard` (0.79 / 0.67 / 0.83), are in the same file. Thresholds belong to the caller: act on a high probability, confirm or escalate on a middle one, hand a low one to a person or a bigger model. The model never refuses.
|
| 173 |
|
| 174 |
## Training
|
| 175 |
|
|
|
|
| 184 |
- Single-hop judgments only. Split a chain of inference into hops.
|
| 185 |
- A choice question is capped at 20 options on `/v1/systemone`; Banking77's 77 options were scored with an extended single-token alphabet for the benchmark only.
|
| 186 |
- English data. Training cut states past 2,048 prompt tokens.
|
| 187 |
+
- Calibration was fitted on the training distribution's calibration split, where HelpSteer2's labels are close to uniform. On traffic skewed like natural HelpSteer2 (over 70% of labels at 3 or 4) score answers are too confident: ECE on Jevals HelpSteer2 is 0.082. Refit on your own labelled traffic before trusting a threshold ([how](eval/RESULTS.md#temperature-fits)).
|
| 188 |
|
| 189 |
## Versioning and license
|
| 190 |
|
|
@@ -16,11 +16,11 @@ The v1 columns are the published [v1](https://huggingface.co/frontier-infra/jeba
|
|
| 16 |
|---|---|---:|---:|---:|---:|---:|---:|---:|
|
| 17 |
| Jevals PubMedQA (300) | noul | 90.3 | 89.7 | 62.0 | 68.6 | 66.8 | 0.025 | 0.0% |
|
| 18 |
| Jevals Banking77 (300, 77 options) | choice | 70.7 | 70.0 | 1.3 | 58.2 | 56.9 | 0.076 | 0.7% |
|
| 19 |
-
| Jevals HelpSteer2 helpfulness (300, 5 levels) | score | 40.7 | 40.3 | 41.7 | 10.
|
| 20 |
-
| Nimble held-out eval (324) | mixed | 81.2 | 78.7 | 17.6 | 65.0 | 65.0 | 0.
|
| 21 |
-
| Kev transfer-v4 test (764) | mixed | 83.8 | 84.0 | 21.5 | 67.9 | 67.0 | 0.
|
| 22 |
-
| Kev decision-v7 test (1,440) | mixed | 81.2 | 80.8 | 20.3 | 71.4 | 70.6 | 0.
|
| 23 |
-
| typed-decisions test (2,000) | mixed | 78.3 | 78.9 | 15.3 |
|
| 24 |
|
| 25 |
Nimble's 13 public human-labelled subsets (3,880 questions), macro accuracy: **77.0** (v1: 77.0).
|
| 26 |
For scale, Bespoke's published table puts Nimble-9B at 74.8 and Jev at 76.0 on the same subsets with their scorer.
|
|
@@ -30,18 +30,18 @@ For scale, Bespoke's published table puts Nimble-9B at 74.8 and Jev at 76.0 on t
|
|
| 30 |
| aegis2 (250) | 78.4 | 79.2 | 38.0 | 35.8 |
|
| 31 |
| boolq (300) | 86.0 | 85.7 | 54.6 | 57.7 |
|
| 32 |
| civil_comments (300) | 86.0 | 86.7 | -7.4 | -4.8 |
|
| 33 |
-
| helpsteer2 (249) | 43.4 | 41.0 |
|
| 34 |
| massive-de-DE (350) | 84.6 | 84.9 | 74.3 | 73.3 |
|
| 35 |
| massive-en-US (350) | 85.4 | 85.7 | 77.3 | 76.0 |
|
| 36 |
| multinli (299) | 84.9 | 87.6 | 67.2 | 70.6 |
|
| 37 |
| paws (250) | 83.6 | 85.6 | 54.0 | 57.2 |
|
| 38 |
| pubmedqa (250) | 72.8 | 68.8 | 38.5 | 34.5 |
|
| 39 |
| squad2 (299) | 79.9 | 80.3 | 42.9 | 43.6 |
|
| 40 |
-
| summeval-consistency (144) | 87.5 | 87.5 |
|
| 41 |
-
| summeval-relevance (240) | 52.5 | 51.7 | 23.
|
| 42 |
| vitaminc-dev (599) | 75.6 | 76.3 | 41.8 | 39.0 |
|
| 43 |
|
| 44 |
-
Where v2 is not better. Jevals HelpSteer2 Decision Score is 10.
|
| 45 |
|
| 46 |
Two things to read carefully, unchanged from v1 because the pool is unchanged. HelpSteer2 and SummEval are not
|
| 47 |
zero-shot for Jeb: the pool trains on the HelpSteer2 **train** split (the Jevals and Nimble items come from
|
|
@@ -50,7 +50,10 @@ are held out entirely), so those rows are held-out items of a seen rubric. And t
|
|
| 50 |
(typed-decisions, Kev decision-v7) are the test splits of sources in the pool.
|
| 51 |
|
| 52 |
Every eval record (per question: option keys, probabilities, pick, label, repeat, option order) is in
|
| 53 |
-
`eval/`, with `eval/results.json` carrying the full metric set. The records were scored on the unmerged adapter
|
|
|
|
|
|
|
|
|
|
| 54 |
the merged weights here reproduce them (see Merge verification below).
|
| 55 |
|
| 56 |
## Robustness to irrelevant content (nonce test)
|
|
@@ -59,7 +62,28 @@ The nonce pass (a fresh UUID planted in the state or appended to the instruction
|
|
| 59 |
|
| 60 |
## Temperature fits
|
| 61 |
|
| 62 |
-
`temperatures.json` carries two fits on the calibration split (689 records, 921 questions: choice 225, noul 147, score 549). The `hard` fit minimises NLL against the argmax label and sharpens (T 0.79 / 0.67 / 0.83). The `train` fit minimises NLL against the training target, the soft or ordinal distribution the model was taught, and softens (T 1.19 / 1.09 / 1.22).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
|
| 64 |
## Merge verification
|
| 65 |
|
|
|
|
| 16 |
|---|---|---:|---:|---:|---:|---:|---:|---:|
|
| 17 |
| Jevals PubMedQA (300) | noul | 90.3 | 89.7 | 62.0 | 68.6 | 66.8 | 0.025 | 0.0% |
|
| 18 |
| Jevals Banking77 (300, 77 options) | choice | 70.7 | 70.0 | 1.3 | 58.2 | 56.9 | 0.076 | 0.7% |
|
| 19 |
+
| Jevals HelpSteer2 helpfulness (300, 5 levels) | score | 40.7 | 40.3 | 41.7 | 10.5 | 11.9 | 0.082 (raw 0.066) | 2.3% |
|
| 20 |
+
| Nimble held-out eval (324) | mixed | 81.2 | 78.7 | 17.6 | 65.0 | 65.0 | 0.070 (raw 0.059) | 0.3% |
|
| 21 |
+
| Kev transfer-v4 test (764) | mixed | 83.8 | 84.0 | 21.5 | 67.9 | 67.0 | 0.027 | 0.3% |
|
| 22 |
+
| Kev decision-v7 test (1,440) | mixed | 81.2 | 80.8 | 20.3 | 71.4 | 70.6 | 0.048 | n/a |
|
| 23 |
+
| typed-decisions test (2,000) | mixed | 78.3 | 78.9 | 15.3 | 57.1 | 56.4 | 0.121 | 0.1% |
|
| 24 |
|
| 25 |
Nimble's 13 public human-labelled subsets (3,880 questions), macro accuracy: **77.0** (v1: 77.0).
|
| 26 |
For scale, Bespoke's published table puts Nimble-9B at 74.8 and Jev at 76.0 on the same subsets with their scorer.
|
|
|
|
| 30 |
| aegis2 (250) | 78.4 | 79.2 | 38.0 | 35.8 |
|
| 31 |
| boolq (300) | 86.0 | 85.7 | 54.6 | 57.7 |
|
| 32 |
| civil_comments (300) | 86.0 | 86.7 | -7.4 | -4.8 |
|
| 33 |
+
| helpsteer2 (249) | 43.4 | 41.0 | 12.3 | 15.2 |
|
| 34 |
| massive-de-DE (350) | 84.6 | 84.9 | 74.3 | 73.3 |
|
| 35 |
| massive-en-US (350) | 85.4 | 85.7 | 77.3 | 76.0 |
|
| 36 |
| multinli (299) | 84.9 | 87.6 | 67.2 | 70.6 |
|
| 37 |
| paws (250) | 83.6 | 85.6 | 54.0 | 57.2 |
|
| 38 |
| pubmedqa (250) | 72.8 | 68.8 | 38.5 | 34.5 |
|
| 39 |
| squad2 (299) | 79.9 | 80.3 | 42.9 | 43.6 |
|
| 40 |
+
| summeval-consistency (144) | 87.5 | 87.5 | 49.4 | 36.0 |
|
| 41 |
+
| summeval-relevance (240) | 52.5 | 51.7 | 23.9 | 25.9 |
|
| 42 |
| vitaminc-dev (599) | 75.6 | 76.3 | 41.8 | 39.0 |
|
| 43 |
|
| 44 |
+
Where v2 is not better. Jevals HelpSteer2 Decision Score is 10.5 against v1's 11.9, and its top level flips over identical repeats on 2.3% of questions against 0.3%. The Nimble public macro is flat (77.0 against 77.0): MultiNLI (84.9 against 87.6) and PAWS (83.6 against 85.6) give back what PubMedQA (72.8 against 68.8) and HelpSteer2 gain. ECE on the Nimble 324 set rises from 0.063 to 0.070. The headline gain is 0.6 points, most of it from the Nimble 324 set (81.2 against 78.7).
|
| 45 |
|
| 46 |
Two things to read carefully, unchanged from v1 because the pool is unchanged. HelpSteer2 and SummEval are not
|
| 47 |
zero-shot for Jeb: the pool trains on the HelpSteer2 **train** split (the Jevals and Nimble items come from
|
|
|
|
| 50 |
(typed-decisions, Kev decision-v7) are the test splits of sources in the pool.
|
| 51 |
|
| 52 |
Every eval record (per question: option keys, probabilities, pick, label, repeat, option order) is in
|
| 53 |
+
`eval/`, with `eval/results.json` carrying the full metric set. The records were scored on the unmerged adapter with the
|
| 54 |
+
previous temperatures (each file's `model.temperatures`); the ECE and Decision Score columns above re-temper those
|
| 55 |
+
same logits to the temperatures now in `temperatures.json` (see Temperature fits), so they differ from `results.json`
|
| 56 |
+
on the sets with score questions. Accuracy is unchanged;
|
| 57 |
the merged weights here reproduce them (see Merge verification below).
|
| 58 |
|
| 59 |
## Robustness to irrelevant content (nonce test)
|
|
|
|
| 62 |
|
| 63 |
## Temperature fits
|
| 64 |
|
| 65 |
+
`temperatures.json` carries two fits on the calibration split (689 records, 921 questions: choice 225, noul 147, score 549), both kept under their names. The `hard` fit minimises NLL against the argmax label and sharpens (T 0.79 / 0.67 / 0.83). The `train` fit minimises NLL against the training target, the soft or ordinal distribution the model was taught, and softens (T 1.19 / 1.09 / 1.22).
|
| 66 |
+
|
| 67 |
+
**Applied since 2026-09-29: `mixed`, choice 1.19 and noul 1.09 from `train`, score 0.83 from `hard`**, the rule the [27B](https://huggingface.co/frontier-infra/jebadiah-27b/blob/main/eval/RESULTS.md#temperature-fits) has applied since 2026-09-26. 9B v2 first shipped with `train` for all three types. The `train` score target is the ordinal kernel, a deliberately smoothed distribution, so fitting T to it softens past what the labels support. No model was retrained and nothing was fitted on an evaluation set: both numbers are the calibration split's own fits. On the calibration split itself, fitting on one half and scoring the other (200 random splits by record family; raw logits for its 549 score questions read through AINode on our fleet, which reproduce the recorded `hard` score fit, 0.837 against 0.833), score ECE is 0.119 with `train`, 0.093 with no temperature and 0.079 with `hard`; NLL 0.953, 0.932, 0.934.
|
| 68 |
+
|
| 69 |
+
Over all 20 evaluation sets (re-tempering the stored per-question logits), question-weighted ECE 0.080 with `train`, 0.079 with no temperature, 0.089 with `hard`, 0.071 applied; macro ECE 0.071, 0.074, 0.093, 0.065; NLL 0.589, 0.591, 0.634, 0.580; macro Decision Score (Jevals) 47.9, 47.8, 46.1, 48.3. Only the sets with score questions change:
|
| 70 |
+
|
| 71 |
+
| Set (score questions present) | ECE, none | ECE, `train` (previous) | ECE, `hard` | ECE, applied | NLL, `train` | NLL, applied |
|
| 72 |
+
|---|---:|---:|---:|---:|---:|---:|
|
| 73 |
+
| Jevals HelpSteer2 (score) | 0.066 | 0.039 | 0.082 | 0.082 | 1.308 | 1.314 |
|
| 74 |
+
| Nimble 324 (mixed) | 0.059 | 0.085 | 0.037 | 0.070 | 0.490 | 0.481 |
|
| 75 |
+
| Kev transfer-v4 (mixed) | 0.037 | 0.025 | 0.067 | 0.027 | 0.442 | 0.440 |
|
| 76 |
+
| Kev decision-v7 (mixed) | 0.059 | 0.050 | 0.087 | 0.048 | 0.530 | 0.534 |
|
| 77 |
+
| typed-decisions (mixed) | 0.125 | 0.156 | 0.078 | 0.121 | 0.591 | 0.553 |
|
| 78 |
+
| Nimble public HelpSteer2 (score) | 0.072 | 0.073 | 0.098 | 0.098 | 1.272 | 1.292 |
|
| 79 |
+
| Nimble public SummEval consistency (score) | 0.149 | 0.195 | 0.118 | 0.118 | 0.532 | 0.462 |
|
| 80 |
+
| Nimble public SummEval relevance (score) | 0.044 | 0.084 | 0.033 | 0.033 | 1.098 | 1.080 |
|
| 81 |
+
|
| 82 |
+
It is not a win everywhere: both HelpSteer2 sets were better calibrated with `train`. We checked whether a separate score temperature per rubric, fitted on the calibration split, would keep HelpSteer2 where it was. It does not, because the split's own HelpSteer2 questions also fit about 0.85: their labels are close to uniform across 0 to 4, while the Jevals and Nimble HelpSteer2 items put over 70% of their labels at 3 or 4, where the model is less accurate and too confident. That is a shift in the label mix, which no fit on this split can see. Per-rubric fits also overfit a split this small (its SummEval part is four articles, and a 17-question rubric's fit ranged from 0.8 to 15 across halves), so one temperature per type stays.
|
| 83 |
+
|
| 84 |
+
Calibration was fitted on the training distribution; refit the temperatures on your own labelled traffic before trusting a threshold. Log the raw distribution (`"calibration": "raw"` on AINode, `--no-temperatures` in the scripts), label a random sample of a few hundred decisions, fit one T per question type against the label on half of them and check ECE and NLL on the other half before you replace the file.
|
| 85 |
+
|
| 86 |
+
ECE here uses 10 equal-width bins on the pick's probability rounded to a whole percent (bin = round(100p) // 10), repeat 0. Binning at floor(10p) gives slightly different values on the flat score sets (on the 9B's Nimble public HelpSteer2 records with `train`, 0.060 against 0.073).
|
| 87 |
|
| 88 |
## Merge verification
|
| 89 |
|
|
@@ -74,7 +74,7 @@
|
|
| 74 |
"temperatures": {
|
| 75 |
"choice": 1.1863,
|
| 76 |
"noul": 1.0903,
|
| 77 |
-
"score":
|
| 78 |
},
|
| 79 |
"note": "same layout, config and tokenizer as the base repo; only the LoRA-targeted language_model weights differ. The served route applies no temperature."
|
| 80 |
}
|
|
|
|
| 74 |
"temperatures": {
|
| 75 |
"choice": 1.1863,
|
| 76 |
"noul": 1.0903,
|
| 77 |
+
"score": 0.8329
|
| 78 |
},
|
| 79 |
"note": "same layout, config and tokenizer as the base repo; only the LoRA-targeted language_model weights differ. The served route applies no temperature."
|
| 80 |
}
|
|
@@ -2,9 +2,25 @@
|
|
| 2 |
"temperatures": {
|
| 3 |
"choice": 1.1863,
|
| 4 |
"noul": 1.0903,
|
| 5 |
-
"score":
|
| 6 |
},
|
| 7 |
-
"applied_target": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
"calib_file": "/workspace/jeb/data-v1/calib.jsonl",
|
| 9 |
"n": {
|
| 10 |
"choice": 225,
|
|
@@ -45,6 +61,26 @@
|
|
| 45 |
"nll_before": 1.1017,
|
| 46 |
"nll_after": 1.0929
|
| 47 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 48 |
}
|
| 49 |
},
|
| 50 |
"nll_before": {
|
|
|
|
| 2 |
"temperatures": {
|
| 3 |
"choice": 1.1863,
|
| 4 |
"noul": 1.0903,
|
| 5 |
+
"score": 0.8329
|
| 6 |
},
|
| 7 |
+
"applied_target": "mixed",
|
| 8 |
+
"applied_fits": {
|
| 9 |
+
"choice": "train",
|
| 10 |
+
"noul": "train",
|
| 11 |
+
"score": "hard"
|
| 12 |
+
},
|
| 13 |
+
"previous": {
|
| 14 |
+
"applied_target": "train",
|
| 15 |
+
"temperatures": {
|
| 16 |
+
"choice": 1.1863,
|
| 17 |
+
"noul": 1.0903,
|
| 18 |
+
"score": 1.2162
|
| 19 |
+
},
|
| 20 |
+
"replaced": "2026-09-29"
|
| 21 |
+
},
|
| 22 |
+
"why": "Refit 2026-09-29 without retraining, the rule the 27B has applied since 2026-09-26. Both fits are unchanged and come from the calibration split only. Score questions calibrate better at the hard fit (T 0.83) than at the train fit (T 1.22), whose ordinal target is deliberately smoothed: on held-out halves of the calibration split score ECE 0.119 -> 0.079. Choice and noul stay on the train fit. Re-tempering the stored logits of the 20 evaluation sets: question-weighted ECE 0.0796 -> 0.0709, macro 0.0709 -> 0.0653, NLL 0.5887 -> 0.5799; accuracy unchanged. Jevals HelpSteer2 gets worse (0.039 -> 0.082). See eval/RESULTS.md, Temperature fits.",
|
| 23 |
+
"top_level_stats_describe": "the train fit (nll_before/nll_after/ece_before/ece_after below are its calibration-split numbers)",
|
| 24 |
"calib_file": "/workspace/jeb/data-v1/calib.jsonl",
|
| 25 |
"n": {
|
| 26 |
"choice": 225,
|
|
|
|
| 61 |
"nll_before": 1.1017,
|
| 62 |
"nll_after": 1.0929
|
| 63 |
}
|
| 64 |
+
},
|
| 65 |
+
"mixed": {
|
| 66 |
+
"choice": {
|
| 67 |
+
"T": 1.1863,
|
| 68 |
+
"nll_before": 0.4594,
|
| 69 |
+
"nll_after": 0.4542,
|
| 70 |
+
"source": "train"
|
| 71 |
+
},
|
| 72 |
+
"noul": {
|
| 73 |
+
"T": 1.0903,
|
| 74 |
+
"nll_before": 0.3776,
|
| 75 |
+
"nll_after": 0.3768,
|
| 76 |
+
"source": "train"
|
| 77 |
+
},
|
| 78 |
+
"score": {
|
| 79 |
+
"T": 0.8329,
|
| 80 |
+
"nll_before": 0.9298,
|
| 81 |
+
"nll_after": 0.9225,
|
| 82 |
+
"source": "hard"
|
| 83 |
+
}
|
| 84 |
}
|
| 85 |
},
|
| 86 |
"nll_before": {
|