Download DIAGNOSTICS.md from vllm-sr/Decision-1.0-Nox-4B: direct link, hf CLI and curl.
- Browser
- Download file 5.68 kB
-
https://huggingface.co/vllm-sr/Decision-1.0-Nox-4B/resolve/e2f752a6d4609cfdab505299127cfdc1e759c1bc/DIAGNOSTICS.md
- Command line
-
hf download hf://vllm-sr/Decision-1.0-Nox-4B@e2f752a6d4609cfdab505299127cfdc1e759c1bc/DIAGNOSTICS.md
-
curl -L -o DIAGNOSTICS.md https://huggingface.co/vllm-sr/Decision-1.0-Nox-4B/resolve/e2f752a6d4609cfdab505299127cfdc1e759c1bc/DIAGNOSTICS.md
Diagnostics
These axes remain separate from headline accuracy. Probability metrics use the shipped distribution; no temperature is fitted on this benchmark. An unavailable full-population metric is shown as —, without replacing it with a successful-only average.
Probability quality on transfer
| Model | Valid / requested | Brier ↓ | NLL ↓ | ECE % ↓ | Coverage at ≤5% error % ↑ | AURC ↓ |
|---|---|---|---|---|---|---|
| Lux | 1046/1046 | 0.3058 | 0.6201 | 5.13 | 57.36 | 0.0639 |
| Nox | 1046/1046 | 0.4169 | 0.9079 | 11.80 | 36.42 | 0.1109 |
| Kev-9B | 1046/1046 | 0.2948 | 0.6028 | 4.45 | 58.03 | 0.0586 |
| Kev-4B | 1046/1046 | 0.3145 | 0.6434 | 2.61 | 56.79 | 0.0659 |
| Qwen3.5-9B | 1046/1046 | 0.3673 | 0.7597 | 9.75 | 47.42 | 0.0872 |
| Decider | 1046/1046 | 0.4122 | 0.7969 | 8.09 | 35.18 | 0.1172 |
| Qwen3.5-4B | 1046/1046 | 0.4199 | 0.8186 | 8.98 | 34.70 | 0.1233 |
| Sol | 1046/1046 | 0.5454 | 1.1094 | 11.25 | 16.16 | 0.2073 |
| Kev-0.8B | 1046/1046 | 0.4839 | 0.9489 | 2.58 | 21.61 | 0.1782 |
| Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
| Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
| Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
| Jev | 1046/1046 | 0.1912 | 0.5743 | 3.35 | 76.96 | 0.0346 |
Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
Option-order robustness
| Model | Valid / requested pairs | Both correct % ↑ | Semantic flip % ↓ | Mean half-L1 ↓ |
|---|---|---|---|---|
| Lux | 36/36 | 77.78 | 8.33 | 0.0923 |
| Nox | 36/36 | 63.89 | 16.67 | 0.0795 |
| Kev-9B | 36/36 | 80.56 | 2.78 | 0.0615 |
| Kev-4B | 36/36 | 77.78 | 5.56 | 0.0667 |
| Qwen3.5-9B | 36/36 | 75.00 | 11.11 | 0.1157 |
| Decider | 36/36 | 83.33 | 11.11 | 0.0701 |
| Qwen3.5-4B | 36/36 | 77.78 | 13.89 | 0.1390 |
| Sol | 36/36 | 41.67 | 38.89 | 0.0694 |
| Kev-0.8B | 36/36 | 55.56 | 16.67 | 0.0808 |
| Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
| Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
| Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
| Jev | 36/36 | 86.11 | 0.00 | 0.0208 |
The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
Missing evidence
| Model | Valid / requested | Intact/control accuracy % ↑ | Mean max P % ↓ | P≥0.9 share % ↓ | Normalized entropy ↑ | Paired confidence drop pp ↑ |
|---|---|---|---|---|---|---|
| Lux | 110/110 | 84.55 | 68.81 | 19.09 | 0.6785 | 19.59 |
| Nox | 110/110 | 72.73 | 78.65 | 27.27 | 0.5072 | 12.64 |
| Kev-9B | 110/110 | 91.82 | 39.61 | 0.00 | 0.9981 | 53.19 |
| Kev-4B | 110/110 | 91.82 | 41.47 | 0.00 | 0.9918 | 51.57 |
| Qwen3.5-9B | 110/110 | 80.91 | 67.62 | 7.27 | 0.7585 | 19.09 |
| Decider | 110/110 | 69.09 | 68.82 | 16.36 | 0.6772 | 13.59 |
| Qwen3.5-4B | 110/110 | 72.73 | 59.89 | 0.91 | 0.8120 | 22.35 |
| Sol | 110/110 | 65.45 | 77.19 | 24.55 | 0.5424 | 5.68 |
| Kev-0.8B | 110/110 | 87.27 | 42.11 | 0.00 | 0.9796 | 40.34 |
| Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
| Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
| Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
| Jev | 110/110 | 93.64 | 62.04 | 16.36 | 0.7657 | 30.30 |
The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
Native contract coverage
| Model | Original probability rows | Transfer probability rows | Transfer truncated questions |
|---|---|---|---|
| Lux | 2720/2720 | 1264/1264 | 0 |
| Nox | 2720/2720 | 1264/1264 | 0 |
| Kev-9B | 2720/2720 | 1264/1264 | 0 |
| Kev-4B | 2720/2720 | 1264/1264 | 0 |
| Qwen3.5-9B | 2720/2720 | 1264/1264 | 0 |
| Decider | 2720/2720 | 1264/1264 | 0 |
| Qwen3.5-4B | 2720/2720 | 1264/1264 | 0 |
| Sol | 2720/2720 | 1264/1264 | 0 |
| Kev-0.8B | 2720/2720 | 1264/1264 | 0 |
| Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
| Laya · English | 2720/2720 | 1264/1264 | 34 |
| Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
| Jev | 2720/2720 | 1264/1264 | — / not observable |
Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation. Server-side truncation for Jev cannot be observed.
Uncertainty
| Model | Overall % | 95% component-bootstrap interval |
|---|---|---|
| Lux | 76.72 | 75.35–78.07 |
| Nox | 73.09 | 71.57–74.56 |
| Kev-9B | 71.89 | 70.42–73.35 |
| Kev-4B | 70.09 | 68.45–71.63 |
| Qwen3.5-9B | 69.73 | 68.27–71.20 |
| Decider | 67.71 | 66.10–69.34 |
| Qwen3.5-4B | 67.29 | 65.89–68.69 |
| Sol | 66.32 | 64.77–67.85 |
| Kev-0.8B | 58.28 | 56.64–59.89 |
| Qwen3.5-2B | 57.24 | 55.74–58.76 |
| Laya · English | 51.03 | 49.43–52.68 |
| Laya · Multilingual | 47.19 | 45.58–48.82 |
| Jev | 81.05 | 79.70–82.35 |
Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.