Jason Brashear PRO
AI & ML interests
Recent Activity
Organizations
Thanks, and a tip: Jeb decides from label logprobs, not chat
Ran it. The 418 HelpSteer2 validation rows that neither Jevals nor Nimble uses want T 1.26 on the 4B and 1.29 on the 9B, with a prompt bootstrap of 1.09 to 1.46 and 1.09 to 1.54. Even reweighted to a flat label mix they want 1.08 and 1.11.
So it lands on the 1.1 side, nowhere near 0.8. The calibration draw is the problem, not the two eval subsets, and the label mix is only part of it.
That T also behaves like our old train T on Jevals and Nimble, and beats the hard T on both sizes. So for v3 the HelpSteer2 part of calibration comes from held-out validation data, not from our own training pool. Thanks for pushing on it.
Thanks again, Dipankar. Good catch on your own 27B row too. We get 0.049 to 0.061 on Nimble public HelpSteer2 as well.
I went back through your SummEval numbers and the mean confidence point against our records. Consistency and relevance, all three sizes, match to the third decimal. So does the confidence gap: at the train T the small models sit within 0.02 of their accuracy on both HelpSteer2 sets, and the hard T pushes that up by 0.06 to 0.07. Your flat mix result matches ours too. The loss roughly halves and never flips.
On your question. The HelpSteer2 part of the calibration split is 161 questions over 87 prompts, and on all of it the 4B gets 0.50 and the 9B gets 0.48.
That number is misleading though. 68 of those 161 are the other four HelpSteer2 attributes (coherence, complexity, verbosity, correctness), which Jevals and Nimble never ask, and coherence and complexity run around 0.7. On helpfulness alone, the question your two eval sets actually ask, it's 0.42 on the 4B and 0.40 on the 9B. Reweighted to a flat mix it's 0.44 and 0.41.
So it isn't well above your 0.43. It's about the same.
You're right about where it comes from. The calibration split is the 5% side of a 95/5 split over the training pool. The HelpSteer2 items in it come from HelpSteer2's train split, drawn evenly across the helpfulness levels. The split is by prompt, so both responses to a prompt land on the same side and no calibration prompt was ever trained on. Jevals and Nimble both sample HelpSteer2's validation split, and the pool builder throws out every validation row.
Here's the part I didn't expect. At the same accuracy, the model is less sure of itself on those calibration items than on the eval items. Fit on helpfulness alone, the calibration slice wants T 0.79 on the 4B and 0.83 on the 9B. A family bootstrap puts the 90% range at 0.64 to 0.99 and 0.67 to 1.04. Fit on Jevals and Nimble directly, the same models want 1.05 to 1.20.
So the calibration fit really is off for HelpSteer2 traffic. It just isn't off because the split is in distribution and flatters the accuracy. Something about that draw makes the model less confident than it is on the validation data, and the label mix is only part of it, same as you found.
The cheap test is sitting right there. HelpSteer2's validation split has 418 rows, over 209 prompts, that neither Jevals nor Nimble uses. They're at the natural label mix and never went near the training pool. We pull raw scores for those from the fleet, fit the score T on them, and see which side it lands on. If it comes out near 1.1, the calibration draw is the problem. If it comes out near 0.8, the two eval subsets are the odd ones. It's a few hundred requests per model, so we'll run it before we touch v3.
What goes on the v3 list: the calibration split draws HelpSteer2 at the natural label mix instead of evenly across levels, a held-out validation slice stays in the eval suite as a standing calibration check, and the calibration split gets bigger than 87 HelpSteer2 prompts.
I really appreciate you pushing on this. :)
Dipankar, thank you for this. You went back to the per-question records and re-tempered them yourself, which is exactly why we put them up, and you caught something real. The 4B and 9B are still shipping the score temperature fit to the training target. The 27B isn't.
Your numbers check out. We rebuilt every one of them from the eval records in the repos.
The mismatch you hit on the Nimble public subsets is binning. Our ECE rounds the confidence to a whole percent before it picks a bin, so round(100c) // 10, and your rows look like floor(10c). On these score sets a lot of the confidence sits right on the bin edges, so it matters more than you'd think. With your bins the 9B Nimble HelpSteer2 comes out 0.060. With ours it's 0.073, and the rows rebuild the stored number exactly. That same difference explains the small gaps in your 4B and 9B columns.
Your 27B row came from our card, so it's in our binning. In yours the 27B Kev decision-v7 goes 0.026 to 0.039 and Jevals HelpSteer2 goes 0.078 to 0.057. None of that changes your conclusion.
On your question, we're taking the 27B rule for the 4B and 9B and eating the HelpSteer2 loss.
I did want to know if splitting score by rubric would buy it back. So we pulled raw logits for all 549 score questions in the calibration split, fit on one half, scored on the other half, and did that 200 times.
It doesn't buy it back. The HelpSteer2 part of the calibration split wants about 0.84 on its own, the same as the pooled fit.
What's going on is the labels. The HelpSteer2 questions in our calibration split are spread pretty evenly across 0 to 4. The Jevals and Nimble HelpSteer2 items are over 70% 3s and 4s. The model is less accurate on that mix and it's too sure of itself there, so it wants to be softer. That's a shift in the label mix, and no rubric split of our calibration data can see it.
Splitting by rubric did help our own typed decision rubrics, which want something closer to 0.5. But the SummEval part of the split is only four articles, the finer per rubric fits swung anywhere from 0.8 to 15 on groups of 17 questions, and a request that hits the server doesn't carry a rubric key anyway. So personally I don't want to ship that off a split this small.
Here's what we're shipping. The 4B and 9B get the hard score fit, 0.83, with choice and noul left alone, same as the 27B. Accuracy doesn't move. The cards will say plainly that HelpSteer2 on both sizes and Kev decision-v7 on the 4B get worse, and why, and that if your traffic looks like HelpSteer2 you should refit on your own labels. Per rubric calibration goes on the list for v3, with a bigger calibration split.
Thanks again. This is exactly the kind of review I was hoping for when we published the records. :)
There are three sizes, 4B, 9B and 27B, all Apache 2.0, with GGUF and MLX builds. Each repo ships a temperatures.json fitted on held-out data, so the probabilities are calibrated per question type.
New this week: pip install jebadiah-decide, then jeb serve. It runs them through Ollama, LM Studio, llama-server, vLLM or MLX and puts /v1/systemone and /v1/decide on localhost. The setup for each runtime is here: https://github.com/getainode/jebadiah/blob/main/docs/run-locally.md
If you try it on a routing or triage step of your own, I'd like to hear how it does.
frontier-infra/jebadiah-open-system-one-decision-models-6ab80765ddd3fa0b3eba5213