File size: 6,327 Bytes
3e46eb1
 
 
 
 
 
 
 
bd45a30
c22a05d
3e46eb1
 
 
 
 
c22a05d
 
3e46eb1
 
c22a05d
3e46eb1
 
3c76737
3e46eb1
 
 
 
 
 
 
bd45a30
c22a05d
3e46eb1
 
 
 
 
c22a05d
 
3e46eb1
 
c22a05d
3e46eb1
 
3c76737
3e46eb1
 
 
 
 
 
 
bd45a30
c22a05d
3e46eb1
 
 
 
 
c22a05d
 
3e46eb1
 
c22a05d
3e46eb1
 
3c76737
3e46eb1
c22a05d
3e46eb1
 
 
 
 
c22a05d
 
3e46eb1
 
 
 
 
c22a05d
 
3e46eb1
 
c22a05d
3e46eb1
 
3c76737
3e46eb1
a3527b0
3e46eb1
 
 
 
 
bd45a30
c22a05d
3e46eb1
 
 
 
 
c22a05d
 
3e46eb1
 
c22a05d
3e46eb1
 
3c76737
3e46eb1
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
# Diagnostics

These axes remain separate from headline accuracy. Probability metrics use the shipped distribution; no temperature is fitted on this benchmark. An unavailable full-population metric is shown as —, without replacing it with a successful-only average.

## Probability quality on transfer

| Model | Valid / requested | Brier ↓ | NLL ↓ | ECE % ↓ | Coverage at ≤5% error % ↑ | AURC ↓ |
|---|---:|---:|---:|---:|---:|---:|
| Lux-9B | 1046/1046 | 0.2911 | 0.5870 | 4.79 | 63.38 | 0.0579 |
| Nox-4B | 1046/1046 | 0.4169 | 0.9079 | 11.80 | 36.42 | 0.1109 |
| Kev-9B | 1046/1046 | 0.2948 | 0.6028 | 4.45 | 58.03 | 0.0586 |
| Kev-4B | 1046/1046 | 0.3145 | 0.6434 | 2.61 | 56.79 | 0.0659 |
| Qwen3.5-9B | 1046/1046 | 0.3673 | 0.7597 | 9.75 | 47.42 | 0.0872 |
| Decider | 1046/1046 | 0.4122 | 0.7969 | 8.09 | 35.18 | 0.1172 |
| Qwen3.5-4B | 1046/1046 | 0.4199 | 0.8186 | 8.98 | 34.70 | 0.1233 |
| Sol-2B | 1046/1046 | 0.5454 | 1.1094 | 11.25 | 16.16 | 0.2073 |
| Eos-0.8B | 1046/1046 | 0.5847 | 1.1406 | 13.96 | 10.23 | 0.2611 |
| Kev-0.8B | 1046/1046 | 0.4839 | 0.9489 | 2.58 | 21.61 | 0.1782 |
| Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
| Kai-0.6B | 1046/1046 | 0.6392 | 1.2391 | 14.36 | 4.02 | 0.3345 |
| Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
| Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
| Jev | 1046/1046 | 0.1912 | 0.5743 | 3.35 | 76.96 | 0.0346 |

Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.

## Option-order robustness

| Model | Valid / requested pairs | Both correct % ↑ | Semantic flip % ↓ | Mean half-L1 ↓ |
|---|---:|---:|---:|---:|
| Lux-9B | 36/36 | 77.78 | 8.33 | 0.0823 |
| Nox-4B | 36/36 | 63.89 | 16.67 | 0.0795 |
| Kev-9B | 36/36 | 80.56 | 2.78 | 0.0615 |
| Kev-4B | 36/36 | 77.78 | 5.56 | 0.0667 |
| Qwen3.5-9B | 36/36 | 75.00 | 11.11 | 0.1157 |
| Decider | 36/36 | 83.33 | 11.11 | 0.0701 |
| Qwen3.5-4B | 36/36 | 77.78 | 13.89 | 0.1390 |
| Sol-2B | 36/36 | 41.67 | 38.89 | 0.0694 |
| Eos-0.8B | 36/36 | 55.56 | 11.11 | 0.1505 |
| Kev-0.8B | 36/36 | 55.56 | 16.67 | 0.0808 |
| Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
| Kai-0.6B | 36/36 | 44.44 | 13.89 | 0.1147 |
| Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
| Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
| Jev | 36/36 | 86.11 | 0.00 | 0.0208 |

The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.

## Missing evidence

| Model | Valid / requested | Intact/control accuracy % ↑ | Mean max P % ↓ | P≥0.9 share % ↓ | Normalized entropy ↑ | Paired confidence drop pp ↑ |
|---|---:|---:|---:|---:|---:|---:|
| Lux-9B | 110/110 | 86.36 | 63.60 | 20.91 | 0.7358 | 26.59 |
| Nox-4B | 110/110 | 72.73 | 78.65 | 27.27 | 0.5072 | 12.64 |
| Kev-9B | 110/110 | 91.82 | 39.61 | 0.00 | 0.9981 | 53.19 |
| Kev-4B | 110/110 | 91.82 | 41.47 | 0.00 | 0.9918 | 51.57 |
| Qwen3.5-9B | 110/110 | 80.91 | 67.62 | 7.27 | 0.7585 | 19.09 |
| Decider | 110/110 | 69.09 | 68.82 | 16.36 | 0.6772 | 13.59 |
| Qwen3.5-4B | 110/110 | 72.73 | 59.89 | 0.91 | 0.8120 | 22.35 |
| Sol-2B | 110/110 | 65.45 | 77.19 | 24.55 | 0.5424 | 5.68 |
| Eos-0.8B | 110/110 | 55.45 | 63.01 | 9.09 | 0.7772 | 1.93 |
| Kev-0.8B | 110/110 | 87.27 | 42.11 | 0.00 | 0.9796 | 40.34 |
| Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
| Kai-0.6B | 110/110 | 37.27 | 52.11 | 18.18 | 0.8263 | 0.62 |
| Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
| Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
| Jev | 110/110 | 93.64 | 62.04 | 16.36 | 0.7657 | 30.30 |

The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. These variants can combine evidence removal, candidate deletion and option permutation. Their confidence shifts are descriptive and do not isolate pure abstention or evidence sensitivity; these are not correctness scores.

## Native contract coverage

| Model | Original probability rows | Transfer probability rows | Transfer truncated questions |
|---|---:|---:|---:|
| Lux-9B | 2720/2720 | 1264/1264 | 0 |
| Nox-4B | 2720/2720 | 1264/1264 | 0 |
| Kev-9B | 2720/2720 | 1264/1264 | 0 |
| Kev-4B | 2720/2720 | 1264/1264 | 0 |
| Qwen3.5-9B | 2720/2720 | 1264/1264 | 0 |
| Decider | 2720/2720 | 1264/1264 | 0 |
| Qwen3.5-4B | 2720/2720 | 1264/1264 | 0 |
| Sol-2B | 2720/2720 | 1264/1264 | 0 |
| Eos-0.8B | 2720/2720 | 1264/1264 | 0 |
| Kev-0.8B | 2720/2720 | 1264/1264 | 0 |
| Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
| Kai-0.6B | 2720/2720 | 1264/1264 | 0 |
| Laya · English | 2720/2720 | 1264/1264 | 34 |
| Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
| Jev | 2720/2720 | 1264/1264 | — / not observable |

Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation. Server-side truncation for Jev cannot be observed.

## Uncertainty

| Model | Overall % | 95% component-bootstrap interval |
|---|---:|---:|
| Lux-9B | 77.40 | 76.01–78.77 |
| Nox-4B | 73.09 | 71.57–74.56 |
| Kev-9B | 71.89 | 70.42–73.35 |
| Kev-4B | 70.09 | 68.45–71.63 |
| Qwen3.5-9B | 69.73 | 68.27–71.20 |
| Decider | 67.71 | 66.10–69.34 |
| Qwen3.5-4B | 67.29 | 65.89–68.69 |
| Sol-2B | 66.32 | 64.77–67.85 |
| Eos-0.8B | 61.89 | 60.26–63.53 |
| Kev-0.8B | 58.28 | 56.64–59.89 |
| Qwen3.5-2B | 57.24 | 55.74–58.76 |
| Kai-0.6B | 53.52 | 51.85–55.25 |
| Laya · English | 51.03 | 49.43–52.68 |
| Laya · Multilingual | 47.19 | 45.58–48.82 |
| Jev | 81.05 | 79.70–82.35 |

Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.