Decision-1.0-Lux-9B / evaluation /DIAGNOSTICS.md
Xunzhuo's picture
Publish clean Decision model repository
cdf4d3e
|
Raw History Blame Contribute Delete
6.33 kB

Diagnostics

These axes remain separate from headline accuracy. Probability metrics use the shipped distribution; no temperature is fitted on this benchmark. An unavailable full-population metric is shown as β€”, without replacing it with a successful-only average.

Probability quality on transfer

Model Valid / requested Brier ↓ NLL ↓ ECE % ↓ Coverage at ≀5% error % ↑ AURC ↓
Lux-9B 1046/1046 0.2911 0.5870 4.79 63.38 0.0579
Nox-4B 1046/1046 0.4169 0.9079 11.80 36.42 0.1109
Kev-9B 1046/1046 0.2948 0.6028 4.45 58.03 0.0586
Kev-4B 1046/1046 0.3145 0.6434 2.61 56.79 0.0659
Qwen3.5-9B 1046/1046 0.3673 0.7597 9.75 47.42 0.0872
Decider 1046/1046 0.4122 0.7969 8.09 35.18 0.1172
Qwen3.5-4B 1046/1046 0.4199 0.8186 8.98 34.70 0.1233
Sol-2B 1046/1046 0.5454 1.1094 11.25 16.16 0.2073
Eos-0.8B 1046/1046 0.5847 1.1406 13.96 10.23 0.2611
Kev-0.8B 1046/1046 0.4839 0.9489 2.58 21.61 0.1782
Qwen3.5-2B 1046/1046 0.5708 1.2327 12.30 11.76 0.2699
Kai-0.6B 1046/1046 0.6392 1.2391 14.36 4.02 0.3345
Laya Β· English 1046/1046 0.5832 1.2275 12.90 0.48 0.2780
Laya Β· Multilingual 1046/1046 0.6873 1.4653 21.36 0.00 0.3460
Jev 1046/1046 0.1912 0.5743 3.35 76.96 0.0346

Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.

Option-order robustness

Model Valid / requested pairs Both correct % ↑ Semantic flip % ↓ Mean half-L1 ↓
Lux-9B 36/36 77.78 8.33 0.0823
Nox-4B 36/36 63.89 16.67 0.0795
Kev-9B 36/36 80.56 2.78 0.0615
Kev-4B 36/36 77.78 5.56 0.0667
Qwen3.5-9B 36/36 75.00 11.11 0.1157
Decider 36/36 83.33 11.11 0.0701
Qwen3.5-4B 36/36 77.78 13.89 0.1390
Sol-2B 36/36 41.67 38.89 0.0694
Eos-0.8B 36/36 55.56 11.11 0.1505
Kev-0.8B 36/36 55.56 16.67 0.0808
Qwen3.5-2B 36/36 55.56 30.56 0.2163
Kai-0.6B 36/36 44.44 13.89 0.1147
Laya Β· English 36/36 44.44 19.44 0.0928
Laya Β· Multilingual 36/36 50.00 22.22 0.1491
Jev 36/36 86.11 0.00 0.0208

The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.

Missing evidence

Model Valid / requested Intact/control accuracy % ↑ Mean max P % ↓ Pβ‰₯0.9 share % ↓ Normalized entropy ↑ Paired confidence drop pp ↑
Lux-9B 110/110 86.36 63.60 20.91 0.7358 26.59
Nox-4B 110/110 72.73 78.65 27.27 0.5072 12.64
Kev-9B 110/110 91.82 39.61 0.00 0.9981 53.19
Kev-4B 110/110 91.82 41.47 0.00 0.9918 51.57
Qwen3.5-9B 110/110 80.91 67.62 7.27 0.7585 19.09
Decider 110/110 69.09 68.82 16.36 0.6772 13.59
Qwen3.5-4B 110/110 72.73 59.89 0.91 0.8120 22.35
Sol-2B 110/110 65.45 77.19 24.55 0.5424 5.68
Eos-0.8B 110/110 55.45 63.01 9.09 0.7772 1.93
Kev-0.8B 110/110 87.27 42.11 0.00 0.9796 40.34
Qwen3.5-2B 110/110 51.82 65.19 9.09 0.7370 4.68
Kai-0.6B 110/110 37.27 52.11 18.18 0.8263 0.62
Laya Β· English 110/110 49.09 67.96 0.00 0.7497 -5.58
Laya Β· Multilingual 110/110 37.27 75.70 29.09 0.5307 -0.49
Jev 110/110 93.64 62.04 16.36 0.7657 30.30

The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. These variants can combine evidence removal, candidate deletion and option permutation. Their confidence shifts are descriptive and do not isolate pure abstention or evidence sensitivity; these are not correctness scores.

Native contract coverage

Model Original probability rows Transfer probability rows Transfer truncated questions
Lux-9B 2720/2720 1264/1264 0
Nox-4B 2720/2720 1264/1264 0
Kev-9B 2720/2720 1264/1264 0
Kev-4B 2720/2720 1264/1264 0
Qwen3.5-9B 2720/2720 1264/1264 0
Decider 2720/2720 1264/1264 0
Qwen3.5-4B 2720/2720 1264/1264 0
Sol-2B 2720/2720 1264/1264 0
Eos-0.8B 2720/2720 1264/1264 0
Kev-0.8B 2720/2720 1264/1264 0
Qwen3.5-2B 2720/2720 1264/1264 0
Kai-0.6B 2720/2720 1264/1264 0
Laya Β· English 2720/2720 1264/1264 34
Laya Β· Multilingual 2720/2720 1264/1264 14
Jev 2720/2720 1264/1264 β€” / not observable

Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation. Server-side truncation for Jev cannot be observed.

Uncertainty

Model Overall % 95% component-bootstrap interval
Lux-9B 77.40 76.01–78.77
Nox-4B 73.09 71.57–74.56
Kev-9B 71.89 70.42–73.35
Kev-4B 70.09 68.45–71.63
Qwen3.5-9B 69.73 68.27–71.20
Decider 67.71 66.10–69.34
Qwen3.5-4B 67.29 65.89–68.69
Sol-2B 66.32 64.77–67.85
Eos-0.8B 61.89 60.26–63.53
Kev-0.8B 58.28 56.64–59.89
Qwen3.5-2B 57.24 55.74–58.76
Kai-0.6B 53.52 51.85–55.25
Laya Β· English 51.03 49.43–52.68
Laya Β· Multilingual 47.19 45.58–48.82
Jev 81.05 79.70–82.35

Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.