Xunzhuo commited on
Commit
e2f752a
·
verified ·
1 Parent(s): 46505c7

Publish qualified Nox Choice semantics and latest Decision comparison

Browse files
DIAGNOSTICS.md CHANGED
@@ -6,9 +6,8 @@ These axes remain separate from headline accuracy. Probability metrics use the s
6
 
7
  | Model | Valid / requested | Brier ↓ | NLL ↓ | ECE % ↓ | Coverage at ≤5% error % ↑ | AURC ↓ |
8
  |---|---:|---:|---:|---:|---:|---:|
9
- | Jev | 1046/1046 | 0.1912 | 0.5743 | 3.35 | 76.96 | 0.0346 |
10
  | Lux | 1046/1046 | 0.3058 | 0.6201 | 5.13 | 57.36 | 0.0639 |
11
- | Nox | 1046/1046 | 0.4350 | 0.9375 | 10.27 | 35.18 | 0.1195 |
12
  | Kev-9B | 1046/1046 | 0.2948 | 0.6028 | 4.45 | 58.03 | 0.0586 |
13
  | Kev-4B | 1046/1046 | 0.3145 | 0.6434 | 2.61 | 56.79 | 0.0659 |
14
  | Qwen3.5-9B | 1046/1046 | 0.3673 | 0.7597 | 9.75 | 47.42 | 0.0872 |
@@ -19,6 +18,7 @@ These axes remain separate from headline accuracy. Probability metrics use the s
19
  | Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
20
  | Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
21
  | Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
 
22
 
23
  Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
24
 
@@ -26,9 +26,8 @@ Coverage at an error threshold keeps whole confidence-tie groups together. These
26
 
27
  | Model | Valid / requested pairs | Both correct % ↑ | Semantic flip % ↓ | Mean half-L1 ↓ |
28
  |---|---:|---:|---:|---:|
29
- | Jev | 36/36 | 86.11 | 0.00 | 0.0208 |
30
  | Lux | 36/36 | 77.78 | 8.33 | 0.0923 |
31
- | Nox | 36/36 | 58.33 | 25.00 | 0.1049 |
32
  | Kev-9B | 36/36 | 80.56 | 2.78 | 0.0615 |
33
  | Kev-4B | 36/36 | 77.78 | 5.56 | 0.0667 |
34
  | Qwen3.5-9B | 36/36 | 75.00 | 11.11 | 0.1157 |
@@ -39,6 +38,7 @@ Coverage at an error threshold keeps whole confidence-tie groups together. These
39
  | Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
40
  | Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
41
  | Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
 
42
 
43
  The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
44
 
@@ -46,7 +46,6 @@ The 36 paired permutations test the same semantics under changed option order. S
46
 
47
  | Model | Valid / requested | Intact/control accuracy % ↑ | Mean max P % ↓ | P≥0.9 share % ↓ | Normalized entropy ↑ | Paired confidence drop pp ↑ |
48
  |---|---:|---:|---:|---:|---:|---:|
49
- | Jev | 110/110 | 93.64 | 62.04 | 16.36 | 0.7657 | 30.30 |
50
  | Lux | 110/110 | 84.55 | 68.81 | 19.09 | 0.6785 | 19.59 |
51
  | Nox | 110/110 | 72.73 | 78.65 | 27.27 | 0.5072 | 12.64 |
52
  | Kev-9B | 110/110 | 91.82 | 39.61 | 0.00 | 0.9981 | 53.19 |
@@ -59,6 +58,7 @@ The 36 paired permutations test the same semantics under changed option order. S
59
  | Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
60
  | Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
61
  | Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
 
62
 
63
  The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
64
 
@@ -66,7 +66,6 @@ The 110 unknowable examples have no scored true class and are excluded from accu
66
 
67
  | Model | Original probability rows | Transfer probability rows | Transfer truncated questions |
68
  |---|---:|---:|---:|
69
- | Jev | 2720/2720 | 1264/1264 | — / not observable |
70
  | Lux | 2720/2720 | 1264/1264 | 0 |
71
  | Nox | 2720/2720 | 1264/1264 | 0 |
72
  | Kev-9B | 2720/2720 | 1264/1264 | 0 |
@@ -79,6 +78,7 @@ The 110 unknowable examples have no scored true class and are excluded from accu
79
  | Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
80
  | Laya · English | 2720/2720 | 1264/1264 | 34 |
81
  | Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
 
82
 
83
  Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation. Server-side truncation for Jev cannot be observed.
84
 
@@ -86,9 +86,8 @@ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and orde
86
 
87
  | Model | Overall % | 95% component-bootstrap interval |
88
  |---|---:|---:|
89
- | Jev | 81.05 | 79.70–82.35 |
90
  | Lux | 76.72 | 75.35–78.07 |
91
- | Nox | 72.84 | 71.33–74.31 |
92
  | Kev-9B | 71.89 | 70.42–73.35 |
93
  | Kev-4B | 70.09 | 68.45–71.63 |
94
  | Qwen3.5-9B | 69.73 | 68.27–71.20 |
@@ -99,5 +98,6 @@ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and orde
99
  | Qwen3.5-2B | 57.24 | 55.74–58.76 |
100
  | Laya · English | 51.03 | 49.43–52.68 |
101
  | Laya · Multilingual | 47.19 | 45.58–48.82 |
 
102
 
103
  Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
 
6
 
7
  | Model | Valid / requested | Brier ↓ | NLL ↓ | ECE % ↓ | Coverage at ≤5% error % ↑ | AURC ↓ |
8
  |---|---:|---:|---:|---:|---:|---:|
 
9
  | Lux | 1046/1046 | 0.3058 | 0.6201 | 5.13 | 57.36 | 0.0639 |
10
+ | Nox | 1046/1046 | 0.4169 | 0.9079 | 11.80 | 36.42 | 0.1109 |
11
  | Kev-9B | 1046/1046 | 0.2948 | 0.6028 | 4.45 | 58.03 | 0.0586 |
12
  | Kev-4B | 1046/1046 | 0.3145 | 0.6434 | 2.61 | 56.79 | 0.0659 |
13
  | Qwen3.5-9B | 1046/1046 | 0.3673 | 0.7597 | 9.75 | 47.42 | 0.0872 |
 
18
  | Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
19
  | Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
20
  | Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
21
+ | Jev | 1046/1046 | 0.1912 | 0.5743 | 3.35 | 76.96 | 0.0346 |
22
 
23
  Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
24
 
 
26
 
27
  | Model | Valid / requested pairs | Both correct % ↑ | Semantic flip % ↓ | Mean half-L1 ↓ |
28
  |---|---:|---:|---:|---:|
 
29
  | Lux | 36/36 | 77.78 | 8.33 | 0.0923 |
30
+ | Nox | 36/36 | 63.89 | 16.67 | 0.0795 |
31
  | Kev-9B | 36/36 | 80.56 | 2.78 | 0.0615 |
32
  | Kev-4B | 36/36 | 77.78 | 5.56 | 0.0667 |
33
  | Qwen3.5-9B | 36/36 | 75.00 | 11.11 | 0.1157 |
 
38
  | Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
39
  | Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
40
  | Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
41
+ | Jev | 36/36 | 86.11 | 0.00 | 0.0208 |
42
 
43
  The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
44
 
 
46
 
47
  | Model | Valid / requested | Intact/control accuracy % ↑ | Mean max P % ↓ | P≥0.9 share % ↓ | Normalized entropy ↑ | Paired confidence drop pp ↑ |
48
  |---|---:|---:|---:|---:|---:|---:|
 
49
  | Lux | 110/110 | 84.55 | 68.81 | 19.09 | 0.6785 | 19.59 |
50
  | Nox | 110/110 | 72.73 | 78.65 | 27.27 | 0.5072 | 12.64 |
51
  | Kev-9B | 110/110 | 91.82 | 39.61 | 0.00 | 0.9981 | 53.19 |
 
58
  | Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
59
  | Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
60
  | Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
61
+ | Jev | 110/110 | 93.64 | 62.04 | 16.36 | 0.7657 | 30.30 |
62
 
63
  The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
64
 
 
66
 
67
  | Model | Original probability rows | Transfer probability rows | Transfer truncated questions |
68
  |---|---:|---:|---:|
 
69
  | Lux | 2720/2720 | 1264/1264 | 0 |
70
  | Nox | 2720/2720 | 1264/1264 | 0 |
71
  | Kev-9B | 2720/2720 | 1264/1264 | 0 |
 
78
  | Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
79
  | Laya · English | 2720/2720 | 1264/1264 | 34 |
80
  | Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
81
+ | Jev | 2720/2720 | 1264/1264 | — / not observable |
82
 
83
  Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation. Server-side truncation for Jev cannot be observed.
84
 
 
86
 
87
  | Model | Overall % | 95% component-bootstrap interval |
88
  |---|---:|---:|
 
89
  | Lux | 76.72 | 75.35–78.07 |
90
+ | Nox | 73.09 | 71.57–74.56 |
91
  | Kev-9B | 71.89 | 70.42–73.35 |
92
  | Kev-4B | 70.09 | 68.45–71.63 |
93
  | Qwen3.5-9B | 69.73 | 68.27–71.20 |
 
98
  | Qwen3.5-2B | 57.24 | 55.74–58.76 |
99
  | Laya · English | 51.03 | 49.43–52.68 |
100
  | Laya · Multilingual | 47.19 | 45.58–48.82 |
101
+ | Jev | 81.05 | 79.70–82.35 |
102
 
103
  Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
EVALUATION.md CHANGED
@@ -4,9 +4,8 @@ The comparison covers **3,766 scored decisions across 54 tasks** and all 13 disp
4
 
5
  | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
6
  |---|---:|---:|---:|---:|---:|---:|---:|
7
- | Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
8
  | Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
9
- | Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** |
10
  | Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
11
  | Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
12
  | Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
@@ -17,8 +16,9 @@ The comparison covers **3,766 scored decisions across 54 tasks** and all 13 disp
17
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
18
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
19
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
 
20
 
21
- Accuracy (%). Rows are sorted by overall score. Bold marks Decision-family cells strictly above every external open or untuned reference in that metric; Jev and the other Decision models are excluded from that threshold.
22
 
23
  ## Scope and weighting
24
 
 
4
 
5
  | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
6
  |---|---:|---:|---:|---:|---:|---:|---:|
7
+ | Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 69.60 | **73.09** |
8
  | Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
 
9
  | Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
10
  | Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
11
  | Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
 
16
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
17
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
18
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
19
+ | Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
20
 
21
+ Accuracy (%). The current model is first, other open models follow by overall score, and Jev is the frontier reference at the end. Bold marks Decision-family cells strictly above every external open or untuned reference in that metric; Jev and the other Decision models are excluded from that threshold.
22
 
23
  ## Scope and weighting
24
 
MATERIALS.json CHANGED
@@ -1,30 +1,26 @@
1
  {
2
- "scope": "13model public comparison; presentation-only roster amendment",
3
- "public_models": 13,
4
- "task_rows": 54,
5
- "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f",
6
- "weights_runtime_temperature_unchanged": true,
7
  "files": {
8
- "DIAGNOSTICS.md": "f0e3a5040f9843bf1480f0383d2f01f7e9a601347d6b1545d68bf861d1f64007",
9
- "EVALUATION.md": "0624b6aebb35e86e8456ed734027ae4590fbf24c91775e4e228fe89454780f14",
10
- "QUESTION-SCALING.md": "898f3554434284ff506dbc63c34859750f07b33dcf71f5b6d4ec7a798a56eb28",
11
- "README.md": "02678d253920bc54d7af864edbee273bb649a5268ede892a6602d6b023f66800",
12
- "SENSITIVITY.md": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048",
13
- "TASKS.md": "ffd8e2da2bed770855e69d54c44479bd66bf27a319684954ad8b55f742706341",
14
- "assets/architecture.png": "8fcfa10b3941ee4aec1f5fba20a25e382a7d5de83074071c9ef68e627996de40",
15
- "assets/architecture.svg": "215d9af6f94cc24847fc1d23a0b2a287b36afdc1e10ea4f5c82895b2fe103002",
16
- "assets/decision-matrix.pdf": "12f963382e810eb2adb8a9fbbc3b3b7e25e75e4c13d56a04db81419e0b565033",
17
- "assets/decision-matrix.png": "29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4",
18
- "assets/decision-matrix.svg": "95e90ed75b9d942e212b621a2968be3c5387d0bf637ca681c885062c5de93456",
19
- "assets/decision-question-scaling.pdf": "d756f60c142ecaf647e59ea5bb3a611a39d1572a6471cd49afb7ef602a23fb3b",
20
- "assets/decision-question-scaling.png": "722e664974b28324693f5f4af6609330c278385ac118ed6f757b8fc0afec64ac",
21
- "assets/decision-question-scaling.svg": "cc591eafa335b5d16edd306d6368679afd283a4b14d5f83fc77780a352b1d315",
22
- "assets/decision-ranking.pdf": "2c5d77b3a2420531e25a35e74389cd74968840117f90ab9f0f79e3e245101a97",
23
- "assets/decision-ranking.png": "a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f",
24
- "assets/decision-ranking.svg": "39be265fee1f5e6a0b051524fadc33786c0a923a28d771a774ea3c01ace460dc",
25
- "assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
26
- "metrics/benchmark.json": "40b1da632cb586a5122fa34f85ea654e2890d6619bbe707b49afe829b4e8d3ce",
27
- "metrics/evaluation-provenance.json": "cab27fc814933826d4a280ab280877c4683524a5ea67773e4e134a00fcf8befa",
28
- "metrics/question-scaling.json": "9407f7ebf86581cdea21e704cdbebb9ced0a485eb980925f48233ebea0f830dc"
29
  }
30
  }
 
1
  {
2
+ "scope": "Latest13model comparison,54tasks and separate diagnostics.",
3
+ "statistics_sha256": "b319a993ac569ea1fa8da027ccb2195868557e4efe232a5d06091db2884086be",
 
 
 
4
  "files": {
5
+ "README.md": "41dfa047d33bdc3aa9feff55928972533ba803a30b332ca380f042268fa1e848",
6
+ "EVALUATION.md": "4159b7939a107e91b13e47345a192eea489f8bd870943e4b397e094b33bba905",
7
+ "TASKS.md": "2fb04bc9c17d805e2abd5e90520df25e79d85aa5296cd42074d4b156f54bf862",
8
+ "DIAGNOSTICS.md": "0b290793565cc5007e9ee7bc79625812c73202c7d8bfce79d8816a0e7a93a965",
9
+ "SENSITIVITY.md": "c3ea31420db4f2e5f17cc7197197a3ddd18abdad773ffb5bf1171b67136c3bc2",
10
+ "WEIGHTING.md": "c3ea31420db4f2e5f17cc7197197a3ddd18abdad773ffb5bf1171b67136c3bc2",
11
+ "metrics/benchmark.json": "178e2e5f58da45f9fbf378ca82b49d7878fc8e9c93eee195a5ef1e7ba0f4fc6b",
12
+ "metrics/evaluation-provenance.json": "a3718fcee42a8cfe0144b01e04346c36a09ac7958154e4561327753d87d5fe8c",
13
+ "assets/decision-ranking.svg": "ab7ca66a9202587dee593714e46be760cb12ac91b6459074bcbbe11c59107d9d",
14
+ "assets/decision-ranking.pdf": "a28d4b0e8649cb7bf2db452a11d93675631af795561726aa63432b05b5888049",
15
+ "assets/decision-ranking.png": "d252593501e8faeb9c5c8cc0a9ae9a96591e7aa8696cf5394cdfec9c0ddac3b7",
16
+ "assets/decision-matrix.svg": "d99102fafe842f2a6bdd4d455f915833d5e0e09e3a22e5211fbf86e5647247ab",
17
+ "assets/decision-matrix.pdf": "5e127c0cd1a6f83b772a3e66ac2e1f4ec19913ac4938371c6f190a23c50a29fe",
18
+ "assets/decision-matrix.png": "c775737aedeecadb0824092d37014a8e2cffa90af0dafcf10d584e0db823ecf8",
19
+ "code/decision_api.py": "6716bdabca3cd2f1aa447d42a6cf53ee62d7ea97984a01f16b4c09fac8aaf17e",
20
+ "bundle-manifest.json": "7d9b06bc25a75b1f131aabd280df0ef2777bf69db98574a300067ace40c61c9c",
21
+ "NULL_DESCRIPTION_RENDERING.json": "3b531cab60cba35648ab71fe4be7fd1af619f02ee337e3beed6296fdfc8b9ec8",
22
+ "model-card-example.json": "1c4fc87d543722e2c3b839cbebd30c5d1370fcaf2d897f97de13079a668ac93e",
23
+ "USAGE.md": "b4c2f44ce120149c7cdee3749283958f34ba9caf4b6492442818c1b6a748c0d3",
24
+ "RUNTIME-RELEASE.json": "325d7e037acab5f61dab4d56d817e7d7e2e873d9b5f0d0df587afd6c33364a46"
 
25
  }
26
  }
NULL_DESCRIPTION_RENDERING.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "rule": "For Choice only, replace a null description with its original key text before tokenization.",
3
+ "preserved": "Non-null descriptions including empty string; keys/order; Noul and Score; original request; weights/tokenizer/temperature/numeric runtime.",
4
+ "source_bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
5
+ "qualification": "Candidate only; full fixed benchmark and independent public loading proof required.",
6
+ "not_a_weight_training_update": true
7
+ }
README.md CHANGED
@@ -32,13 +32,12 @@ tags:
32
 
33
  ## Measured capability
34
 
35
- **72.84% overall accuracy** across 3,766 scored decisions and 54 tasks. Nox leads Kev-4B by **2.75 percentage points** on this decision-focused comparison; reading and transfer remain opportunities to improve.
36
 
37
  | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
38
  |---|---:|---:|---:|---:|---:|---:|---:|
39
- | Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
40
  | Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
41
- | Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** |
42
  | Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
43
  | Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
44
  | Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
@@ -49,6 +48,7 @@ tags:
49
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
50
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
51
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
 
52
 
53
  Accuracy (%). Overall weights: Decisions **30%**, Composition **25%**, Reading **15%**, Inference **15%**, Transfer **15%**. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.
54
 
@@ -78,6 +78,8 @@ print(model.decide(**REQUEST)["answers"])
78
 
79
  [Tested SystemOne request and output](model-card-example.json) · [Install and typed API guide](USAGE.md)
80
 
 
 
81
  The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.
82
 
83
  ## Architecture
 
32
 
33
  ## Measured capability
34
 
35
+ **73.09% overall accuracy** across 3,766 scored decisions and 54 tasks. Nox leads Kev-4B by **2.99 percentage points** on this decision-focused comparison; reading and transfer remain opportunities to improve.
36
 
37
  | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
38
  |---|---:|---:|---:|---:|---:|---:|---:|
39
+ | Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 69.60 | **73.09** |
40
  | Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
 
41
  | Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
42
  | Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
43
  | Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
 
48
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
49
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
50
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
51
+ | Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
52
 
53
  Accuracy (%). Overall weights: Decisions **30%**, Composition **25%**, Reading **15%**, Inference **15%**, Transfer **15%**. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.
54
 
 
78
 
79
  [Tested SystemOne request and output](model-card-example.json) · [Install and typed API guide](USAGE.md)
80
 
81
+ Choice candidates with a null description use their ID text, which may increase input tokens.
82
+
83
  The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.
84
 
85
  ## Architecture
RUNTIME-RELEASE.json CHANGED
@@ -1,30 +1,11 @@
1
  {
2
- "format": "decision-runtime-patch-v1",
3
- "family": "Nox",
4
- "release_tag": "v1.3.1",
5
- "change_kind": "runtime_only",
6
- "weights_revision": "ad089ad3a5dc9a7a21e6d96db546bb53e2212654",
7
- "source_bundle_manifest_sha256": "92d7f5be5e1ef21ee682f574b35cc01de4dd8ab16b8aa6edf0dfbdfcfcfabba3",
8
- "bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
9
- "changed_inference_files": [
10
- "code/decision_api.py"
11
- ],
12
- "weights_tokenizer_prompts_calibration_profile_unchanged": true,
13
- "default_public_parity": {
14
- "requests": 2856,
15
- "answers": 3160,
16
- "fixtures": 58,
17
- "raw_logits_probabilities_typed_outputs_exact": true,
18
- "receipt_sha256": "946dba5c31074b63efa3223ede17857e5f9d1b7b3b44de08cd9add9ceece878c"
19
- },
20
- "timing": {
21
- "receipt_sha256": "501f25efd1813b566d77e98278229440b7a3b4a2358def1462c0171a33e9f3be",
22
- "shape": "distinct fixed-length questions; Q=32; 499 tokens per question",
23
- "reference_ms": 419.124,
24
- "candidate_ms": 414.199,
25
- "same_measured_api_sha256": "273f6f10f22d5a68b8db34cfcbd35407fb43d8030f6d7cd188bcf118dd90a152",
26
- "no_shared_state_neural_cache_claim": true
27
- },
28
- "capability_scores_unchanged": true,
29
- "downloaded_package_offline_proof_required_before_promotion": true
30
  }
 
1
  {
2
+ "release": "v1.3.2",
3
+ "kind": "SystemOne Choice null-description semantics",
4
+ "rule": "A null Choice description uses the original candidate key as its description.",
5
+ "weights_tokenizer_temperature_unchanged": true,
6
+ "explicit_descriptions_and_Noul_Score_unchanged": true,
7
+ "bundle_manifest_sha256": "7d9b06bc25a75b1f131aabd280df0ef2777bf69db98574a300067ace40c61c9c",
8
+ "qualified_statistics_sha256": "b319a993ac569ea1fa8da027ccb2195868557e4efe232a5d06091db2884086be",
9
+ "exact_public_contract_proof_sha256": "ebf27c7aa2bb4a12ba2cbb27e4096e46cb43ac4a20b9d3e1bc34f62777a1e63a",
10
+ "runtime_note": "This is an API rendering improvement, not a newly trained checkpoint."
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
11
  }
SENSITIVITY.md CHANGED
@@ -1,10 +1,9 @@
1
- Product weights were changed after earlier results were observed. These are identical model predictions under different weights, not trained-model improvements.
2
 
3
  | Model | Current 30/25/15/15/15 | Earlier 25/25/15/15/20 | Original four-panel mean |
4
  |---|---:|---:|---:|
5
- | Jev | 81.05 | 81.45 | 82.45 |
6
  | Lux | 76.72 | 76.43 | 79.06 |
7
- | Nox | 72.84 | 72.09 | 75.03 |
8
  | Kev-9B | 71.89 | 72.01 | 73.19 |
9
  | Kev-4B | 70.09 | 70.30 | 71.73 |
10
  | Qwen3.5-9B | 69.73 | 69.70 | 71.99 |
@@ -15,3 +14,4 @@ Product weights were changed after earlier results were observed. These are iden
15
  | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
16
  | Laya · English | 51.03 | 50.85 | 51.76 |
17
  | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
 
 
1
+ The same latest model predictions are shown under three weighting schemes. Product weights were chosen after earlier results were observed; this sensitivity view is not a trained-model improvement.
2
 
3
  | Model | Current 30/25/15/15/15 | Earlier 25/25/15/15/20 | Original four-panel mean |
4
  |---|---:|---:|---:|
 
5
  | Lux | 76.72 | 76.43 | 79.06 |
6
+ | Nox | 73.09 | 72.42 | 75.03 |
7
  | Kev-9B | 71.89 | 72.01 | 73.19 |
8
  | Kev-4B | 70.09 | 70.30 | 71.73 |
9
  | Qwen3.5-9B | 69.73 | 69.70 | 71.99 |
 
14
  | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
15
  | Laya · English | 51.03 | 50.85 | 51.76 |
16
  | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
17
+ | Jev | 81.05 | 81.45 | 82.45 |
TASKS.md CHANGED
@@ -5,93 +5,93 @@ Accuracy (%) on the same requested rows. Bold marks a Decision-family result str
5
  <details>
6
  <summary>Decisions · 10 tasks</summary>
7
 
8
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
9
- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
10
- | News classification | 128 | 85.16 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 |
11
- | Boolean constraints | 64 | 100.00 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 |
12
- | Entity classification | 112 | 96.43 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 |
13
- | Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 |
14
- | Evidence placement | 96 | 30.21 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 |
15
- | Ordered rubric | 64 | 100.00 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 |
16
- | Relation composition | 96 | 56.25 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 |
17
- | Scoped evidence | 96 | 89.58 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 |
18
- | State tracking | 96 | 33.33 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 |
19
- | In / out of menu | 64 | 100.00 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 |
20
 
21
  </details>
22
 
23
  <details>
24
  <summary>Composition · 10 tasks</summary>
25
 
26
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
27
- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
28
- | Record identity | 80 | 76.25 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 |
29
- | Capacity assignment | 80 | 68.75 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 |
30
- | Constraint assignment | 80 | 55.00 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 |
31
- | Intent routing · EN | 120 | 91.67 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 |
32
- | Intent routing · ZH | 120 | 88.33 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 |
33
- | Multiset reconciliation | 80 | 58.75 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 |
34
- | Ordered service loss | 80 | 42.50 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 |
35
- | Conflicting rule closure | 80 | 66.25 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 |
36
- | Temporal exclusion | 80 | 41.25 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 |
37
- | Transaction recovery | 80 | 75.00 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 |
38
 
39
  </details>
40
 
41
  <details>
42
  <summary>Reading · 3 tasks</summary>
43
 
44
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
45
- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
46
- | Yes / no reading | 160 | 92.50 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 |
47
- | Reading · EN | 160 | 96.88 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 |
48
- | Reading · ZH | 160 | 96.25 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 |
49
 
50
  </details>
51
 
52
  <details>
53
  <summary>Inference · 4 tasks</summary>
54
 
55
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
56
- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
57
- | Contextual reasoning | 120 | 86.67 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 |
58
- | Answerability | 120 | 90.83 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 |
59
- | Textual entailment | 120 | 82.50 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 |
60
- | Scientific inference | 120 | 99.17 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 |
61
 
62
  </details>
63
 
64
  <details>
65
  <summary>Transfer · 27 tasks</summary>
66
 
67
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
68
- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
69
- | Buried emotion | 20 | 70.00 | 40.00 | 20.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 |
70
- | Buried paraphrase | 20 | 95.00 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 |
71
- | Buried entailment | 20 | 95.00 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 |
72
- | Buried offensive-language detection | 20 | 55.00 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 |
73
- | Combined policy conditions | 32 | 96.88 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 |
74
- | Policy exceptions | 32 | 100.00 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 |
75
- | Policy negation | 32 | 90.62 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 |
76
- | Authorization contrast | 40 | 100.00 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 |
77
- | Deadline contrast | 40 | 92.50 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 |
78
- | Emotion | 80 | 67.50 | 61.25 | 40.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 |
79
- | MMLU | 80 | 88.75 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 |
80
- | MMLU-Pro | 200 | 84.00 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 |
81
- | Paraphrase | 80 | 87.50 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 |
82
- | Question entailment | 80 | 91.25 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 |
83
- | Science questions | 80 | 100.00 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 |
84
- | Offensive-language detection | 80 | 76.25 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 |
85
- | Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 |
86
- | Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 |
87
- | Evidence control · deadline | 10 | 100.00 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 |
88
- | Evidence control · late fee | 10 | 100.00 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 |
89
- | Evidence control · quantity limit | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 |
90
- | Evidence control · return window | 10 | 50.00 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 |
91
- | Evidence control · shipping delay | 10 | 90.00 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 |
92
- | Evidence control · sla response | 10 | 100.00 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 |
93
- | Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 |
94
- | Evidence control · volume discount | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 |
95
- | Evidence control · warranty claim | 10 | 90.00 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 |
96
 
97
  </details>
 
5
  <details>
6
  <summary>Decisions · 10 tasks</summary>
7
 
8
+ | Task | n | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Jev |
9
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
10
+ | News classification | 128 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 | 85.16 |
11
+ | Boolean constraints | 64 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 | 100.00 |
12
+ | Entity classification | 112 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 | 96.43 |
13
+ | Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 | 100.00 |
14
+ | Evidence placement | 96 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 | 30.21 |
15
+ | Ordered rubric | 64 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 | 100.00 |
16
+ | Relation composition | 96 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 | 56.25 |
17
+ | Scoped evidence | 96 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 | 89.58 |
18
+ | State tracking | 96 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 | 33.33 |
19
+ | In / out of menu | 64 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 | 100.00 |
20
 
21
  </details>
22
 
23
  <details>
24
  <summary>Composition · 10 tasks</summary>
25
 
26
+ | Task | n | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Jev |
27
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
28
+ | Record identity | 80 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 | 76.25 |
29
+ | Capacity assignment | 80 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 | 68.75 |
30
+ | Constraint assignment | 80 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 | 55.00 |
31
+ | Intent routing · EN | 120 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 | 91.67 |
32
+ | Intent routing · ZH | 120 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 | 88.33 |
33
+ | Multiset reconciliation | 80 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 | 58.75 |
34
+ | Ordered service loss | 80 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 | 42.50 |
35
+ | Conflicting rule closure | 80 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 | 66.25 |
36
+ | Temporal exclusion | 80 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 | 41.25 |
37
+ | Transaction recovery | 80 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 | 75.00 |
38
 
39
  </details>
40
 
41
  <details>
42
  <summary>Reading · 3 tasks</summary>
43
 
44
+ | Task | n | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Jev |
45
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
46
+ | Yes / no reading | 160 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 | 92.50 |
47
+ | Reading · EN | 160 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 | 96.88 |
48
+ | Reading · ZH | 160 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 | 96.25 |
49
 
50
  </details>
51
 
52
  <details>
53
  <summary>Inference · 4 tasks</summary>
54
 
55
+ | Task | n | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Jev |
56
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
57
+ | Contextual reasoning | 120 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 | 86.67 |
58
+ | Answerability | 120 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 | 90.83 |
59
+ | Textual entailment | 120 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 | 82.50 |
60
+ | Scientific inference | 120 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 | 99.17 |
61
 
62
  </details>
63
 
64
  <details>
65
  <summary>Transfer · 27 tasks</summary>
66
 
67
+ | Task | n | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Jev |
68
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
69
+ | Buried emotion | 20 | 40.00 | 25.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 | 70.00 |
70
+ | Buried paraphrase | 20 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 | 95.00 |
71
+ | Buried entailment | 20 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 | 95.00 |
72
+ | Buried offensive-language detection | 20 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 | 55.00 |
73
+ | Combined policy conditions | 32 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 | 96.88 |
74
+ | Policy exceptions | 32 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 | 100.00 |
75
+ | Policy negation | 32 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 | 90.62 |
76
+ | Authorization contrast | 40 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 | 100.00 |
77
+ | Deadline contrast | 40 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 | 92.50 |
78
+ | Emotion | 80 | 61.25 | 60.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 | 67.50 |
79
+ | MMLU | 80 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 | 88.75 |
80
+ | MMLU-Pro | 200 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 | 84.00 |
81
+ | Paraphrase | 80 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 | 87.50 |
82
+ | Question entailment | 80 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 | 91.25 |
83
+ | Science questions | 80 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 | 100.00 |
84
+ | Offensive-language detection | 80 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 | 76.25 |
85
+ | Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 | 100.00 |
86
+ | Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 | 100.00 |
87
+ | Evidence control · deadline | 10 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 | 100.00 |
88
+ | Evidence control · late fee | 10 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 | 100.00 |
89
+ | Evidence control · quantity limit | 10 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 | 100.00 |
90
+ | Evidence control · return window | 10 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 | 50.00 |
91
+ | Evidence control · shipping delay | 10 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 | 90.00 |
92
+ | Evidence control · sla response | 10 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 | 100.00 |
93
+ | Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 | 100.00 |
94
+ | Evidence control · volume discount | 10 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 | 100.00 |
95
+ | Evidence control · warranty claim | 10 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 | 90.00 |
96
 
97
  </details>
USAGE.md CHANGED
@@ -49,12 +49,14 @@ The model card must publish output from that model's real run; this document inv
49
 
50
  | Type | Input criteria | Answer |
51
  |---|---|---|
52
- | Choice | Ordered mapping of 2–255 external IDs to complete descriptions | `probabilities`, selected `choice` ID, `confidence` |
53
  | Noul | Optional mapping containing only `false` / `true` descriptions | `noul`: P(true); a hard judgment uses `>= 0.5` |
54
  | Score | Ordered list of 2–10 rubric descriptions | `probabilities` over string indices, expected level index `score`, `legend`, `confidence` |
55
 
56
  Response shape is `{"model": name, "answers": {question_name: answer}, "usage": {"input_tokens": total, "scored_questions": count}}`. Choice ties select the earliest candidate in insertion order. Score returns the expected ordinal **index**, not an arbitrary supplied numeric value; this adapter does not implement a supplied-values extension. Noul 0.5 is interpreted as true. Confidence is `(K * max(p) - 1)/(K - 1)`, clipped to [0,1], and is not a claimed reproduction of Jev's confidence statistic. Shipped temperature calibration does not make every confidence value a correctness guarantee.
57
 
 
 
58
  Question names are preserved as opaque bookkeeping IDs; candidate IDs and descriptions use the frozen renderer. Native objects are deterministically serialized with sorted JSON keys; strings preserve their contents. Do not interpret object/string field-order differences as identical token inputs. Non-finite or non-JSON inputs are rejected.
59
 
60
  The bound is **16,384 tokens per complete question**, including its state, instructions, all candidates and readout suffix. Questions are separate sequences, grouped in fixed batches of eight; the state is repeated for each question and counted repeatedly in `usage.input_tokens`. All questions are encoded before any forward pass. If any exceeds the bound, the whole call raises `ValueError` without truncation or partial answers. A lower `max_length` can be chosen at load time; a higher limit is rejected. This native wrapper does not impose the Studio's separate 16-question UI limit.
 
49
 
50
  | Type | Input criteria | Answer |
51
  |---|---|---|
52
+ | Choice | Ordered mapping of 2–255 external IDs to descriptions or null | `probabilities`, selected `choice` ID, `confidence` |
53
  | Noul | Optional mapping containing only `false` / `true` descriptions | `noul`: P(true); a hard judgment uses `>= 0.5` |
54
  | Score | Ordered list of 2–10 rubric descriptions | `probabilities` over string indices, expected level index `score`, `legend`, `confidence` |
55
 
56
  Response shape is `{"model": name, "answers": {question_name: answer}, "usage": {"input_tokens": total, "scored_questions": count}}`. Choice ties select the earliest candidate in insertion order. Score returns the expected ordinal **index**, not an arbitrary supplied numeric value; this adapter does not implement a supplied-values extension. Noul 0.5 is interpreted as true. Confidence is `(K * max(p) - 1)/(K - 1)`, clipped to [0,1], and is not a claimed reproduction of Jev's confidence statistic. Shipped temperature calibration does not make every confidence value a correctness guarantee.
57
 
58
+ For Choice, a null description uses the original candidate ID as its semantic description. Explicit descriptions, including an empty string, are preserved. The original request object is not mutated. Noul and Score are unchanged.
59
+
60
  Question names are preserved as opaque bookkeeping IDs; candidate IDs and descriptions use the frozen renderer. Native objects are deterministically serialized with sorted JSON keys; strings preserve their contents. Do not interpret object/string field-order differences as identical token inputs. Non-finite or non-JSON inputs are rejected.
61
 
62
  The bound is **16,384 tokens per complete question**, including its state, instructions, all candidates and readout suffix. Questions are separate sequences, grouped in fixed batches of eight; the state is repeated for each question and counted repeatedly in `usage.input_tokens`. All questions are encoded before any forward pass. If any exceeds the bound, the whole call raises `ValueError` without truncation or partial answers. A lower `max_length` can be chosen at load time; a higher limit is rejected. This native wrapper does not impose the Studio's separate 16-question UI limit.
WEIGHTING.md ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ The same latest model predictions are shown under three weighting schemes. Product weights were chosen after earlier results were observed; this sensitivity view is not a trained-model improvement.
2
+
3
+ | Model | Current 30/25/15/15/15 | Earlier 25/25/15/15/20 | Original four-panel mean |
4
+ |---|---:|---:|---:|
5
+ | Lux | 76.72 | 76.43 | 79.06 |
6
+ | Nox | 73.09 | 72.42 | 75.03 |
7
+ | Kev-9B | 71.89 | 72.01 | 73.19 |
8
+ | Kev-4B | 70.09 | 70.30 | 71.73 |
9
+ | Qwen3.5-9B | 69.73 | 69.70 | 71.99 |
10
+ | Decider | 67.71 | 67.97 | 71.75 |
11
+ | Qwen3.5-4B | 67.29 | 67.24 | 70.25 |
12
+ | Sol | 66.32 | 65.48 | 70.14 |
13
+ | Kev-0.8B | 58.28 | 58.33 | 59.75 |
14
+ | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
15
+ | Laya · English | 51.03 | 50.85 | 51.76 |
16
+ | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
17
+ | Jev | 81.05 | 81.45 | 82.45 |
assets/decision-matrix.pdf CHANGED
Binary files a/assets/decision-matrix.pdf and b/assets/decision-matrix.pdf differ
 
assets/decision-matrix.png CHANGED

Git LFS Details

  • SHA256: 29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4
  • Pointer size: 131 Bytes
  • Size of remote file: 325 kB

Git LFS Details

  • SHA256: c775737aedeecadb0824092d37014a8e2cffa90af0dafcf10d584e0db823ecf8
  • Pointer size: 131 Bytes
  • Size of remote file: 326 kB
assets/decision-matrix.svg CHANGED
assets/decision-ranking.pdf CHANGED
Binary files a/assets/decision-ranking.pdf and b/assets/decision-ranking.pdf differ
 
assets/decision-ranking.png CHANGED

Git LFS Details

  • SHA256: a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f
  • Pointer size: 131 Bytes
  • Size of remote file: 215 kB

Git LFS Details

  • SHA256: d252593501e8faeb9c5c8cc0a9ae9a96591e7aa8696cf5394cdfec9c0ddac3b7
  • Pointer size: 131 Bytes
  • Size of remote file: 215 kB
assets/decision-ranking.svg CHANGED
bundle-manifest.json CHANGED
@@ -1,12 +1,17 @@
1
  {
2
  "format": "research-pointer-bundle-v1",
3
- "status": "joint-serving-candidate-awaiting-default-public-parity",
4
  "files": [
5
  {
6
  "file": "NORMALIZATION_RUNTIME.md",
7
  "bytes": 756,
8
  "sha256": "cf8e6ce1f07687a68b6adeb98e6704b8f292cb6adc7b10e4b84e8e70aa5472e5"
9
  },
 
 
 
 
 
10
  {
11
  "file": "RUNTIME_BINDING.json",
12
  "bytes": 2621,
@@ -54,8 +59,8 @@
54
  },
55
  {
56
  "file": "code/decision_api.py",
57
- "bytes": 10952,
58
- "sha256": "273f6f10f22d5a68b8db34cfcbd35407fb43d8030f6d7cd188bcf118dd90a152"
59
  },
60
  {
61
  "file": "code/decision_model.py",
@@ -221,7 +226,7 @@
221
  }
222
  ],
223
  "source_model_code_sha256": "d3e28489c09f3bd7130e2d43d92e0b5c4a08e25b09b21303904defb0ff1c3646",
224
- "source_api_code_sha256": "273f6f10f22d5a68b8db34cfcbd35407fb43d8030f6d7cd188bcf118dd90a152",
225
  "dev_sha256": "45b4cd46acc4be53b95c7ed8f333b3533972296a4d0557c7b8c48c9cab6ced61",
226
  "production_predictions_sha256": "ba311f5a6bc73182651bdf5c9744f0ed95921110501ddc93dcfd20a1d06bceff",
227
  "temperature_sha256": "69d80e5b215e2c1e6872f146fdb7ea5aa95c9fd678ef96208960d50e4e7b635d",
@@ -230,7 +235,7 @@
230
  "input_length_limit": 16384,
231
  "original_checkpoint_name": "winner",
232
  "no_publication_performed": true,
233
- "source_bundle_manifest_sha256": "d437bc0149bcb8c9891fbb33c5abc4c2336c3a89ab6cfa0981da5a7f1d19f1f4",
234
  "normalization_profile_sha256": "be32858d15233e0a3fbee0e4257fb02be0b3439deee4eb9c3f61151df7b73850",
235
  "public_wrapper_included": true,
236
  "source_published_bundle_manifest_sha256": "92d7f5be5e1ef21ee682f574b35cc01de4dd8ab16b8aa6edf0dfbdfcfcfabba3",
 
1
  {
2
  "format": "research-pointer-bundle-v1",
3
+ "status": "null-description-candidate-awaiting-public-proof-and-full-regression",
4
  "files": [
5
  {
6
  "file": "NORMALIZATION_RUNTIME.md",
7
  "bytes": 756,
8
  "sha256": "cf8e6ce1f07687a68b6adeb98e6704b8f292cb6adc7b10e4b84e8e70aa5472e5"
9
  },
10
+ {
11
+ "file": "NULL_DESCRIPTION_RENDERING.json",
12
+ "bytes": 514,
13
+ "sha256": "3b531cab60cba35648ab71fe4be7fd1af619f02ee337e3beed6296fdfc8b9ec8"
14
+ },
15
  {
16
  "file": "RUNTIME_BINDING.json",
17
  "bytes": 2621,
 
59
  },
60
  {
61
  "file": "code/decision_api.py",
62
+ "bytes": 10978,
63
+ "sha256": "6716bdabca3cd2f1aa447d42a6cf53ee62d7ea97984a01f16b4c09fac8aaf17e"
64
  },
65
  {
66
  "file": "code/decision_model.py",
 
226
  }
227
  ],
228
  "source_model_code_sha256": "d3e28489c09f3bd7130e2d43d92e0b5c4a08e25b09b21303904defb0ff1c3646",
229
+ "source_api_code_sha256": "6716bdabca3cd2f1aa447d42a6cf53ee62d7ea97984a01f16b4c09fac8aaf17e",
230
  "dev_sha256": "45b4cd46acc4be53b95c7ed8f333b3533972296a4d0557c7b8c48c9cab6ced61",
231
  "production_predictions_sha256": "ba311f5a6bc73182651bdf5c9744f0ed95921110501ddc93dcfd20a1d06bceff",
232
  "temperature_sha256": "69d80e5b215e2c1e6872f146fdb7ea5aa95c9fd678ef96208960d50e4e7b635d",
 
235
  "input_length_limit": 16384,
236
  "original_checkpoint_name": "winner",
237
  "no_publication_performed": true,
238
+ "source_bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
239
  "normalization_profile_sha256": "be32858d15233e0a3fbee0e4257fb02be0b3439deee4eb9c3f61151df7b73850",
240
  "public_wrapper_included": true,
241
  "source_published_bundle_manifest_sha256": "92d7f5be5e1ef21ee682f574b35cc01de4dd8ab16b8aa6edf0dfbdfcfcfabba3",
code/decision_api.py CHANGED
@@ -55,7 +55,7 @@ def question_row(state, name, question):
55
  if not isinstance(criteria,dict) or not 2<=len(criteria)<=255:
56
  raise ValueError('choice requires a mapping of 2..255 criteria')
57
  if not all(isinstance(k,str) for k in criteria):raise ValueError('Choice keys must be strings')
58
- options=[{'key':key,'description':value} for key,value in criteria.items()]
59
  # The question name is used for bookkeeping only; encoders never render id.
60
  return {'id':name,'state':state,'instructions':question['instructions'],
61
  'options':options,'task_type':kind,'family':'inference'}
 
55
  if not isinstance(criteria,dict) or not 2<=len(criteria)<=255:
56
  raise ValueError('choice requires a mapping of 2..255 criteria')
57
  if not all(isinstance(k,str) for k in criteria):raise ValueError('Choice keys must be strings')
58
+ options=[{'key':key,'description':key if value is None else value} for key,value in criteria.items()]
59
  # The question name is used for bookkeeping only; encoders never render id.
60
  return {'id':name,'state':state,'instructions':question['instructions'],
61
  'options':options,'task_type':kind,'family':'inference'}
metrics/benchmark.json CHANGED
@@ -9,7 +9,7 @@
9
  },
10
  "metric_design": "User-requested, outcome-informed decision-priority weights; observed regression data, not blind testing.",
11
  "protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
12
- "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
13
  "models": {
14
  "Jev": {
15
  "panels": {
@@ -10330,11 +10330,11 @@
10330
  ]
10331
  },
10332
  "transfer_v9_test": {
10333
- "point": 0.6797323135755258,
10334
- "exact_fraction": "711/1046",
10335
  "ci95": [
10336
- 0.6488257294965968,
10337
- 0.7102012834523006
10338
  ]
10339
  },
10340
  "original_four_panel_mean": {
@@ -10346,11 +10346,11 @@
10346
  ]
10347
  },
10348
  "weighted_mean": {
10349
- "point": 0.728414460131567,
10350
- "exact_fraction": "102402253/140582400",
10351
  "ci95": [
10352
- 0.7132642380063722,
10353
- 0.7430890411518838
10354
  ]
10355
  }
10356
  },
@@ -10496,50 +10496,50 @@
10496
  "clean": {
10497
  "requested": 1046,
10498
  "valid_probability_rows": 1046,
10499
- "correct": 711,
10500
- "accuracy": 0.6797323135755258,
10501
  "full_probability_metrics": {
10502
  "n": 1046,
10503
- "nll": 0.9375416500577584,
10504
- "acc": 0.6797323135755258,
10505
- "ece": 0.10273742120233656,
10506
- "brier": 0.4350401558940114,
10507
- "mean_conf": 0.7714702411023553,
10508
- "confident_error_rate": 0.05449330783938815,
10509
- "coverage_at_0_9": 0.5239005736137667,
10510
- "accuracy_at_0_9": 0.8959854014598541,
10511
- "coverage_at_5pct_error": 0.35181644359464626,
10512
- "coverage_at_1pct_error": 0.14149139579349904,
10513
- "aurc": 0.11948583282589058,
10514
- "error_rate_at_0_9": 0.10401459854014598,
10515
- "confidence_bias": 0.09173792752682952,
10516
  "top_bins": {
10517
  "0.9": {
10518
- "n": 548,
10519
- "errors": 57,
10520
- "error_rate": 0.10401459854014598
10521
  },
10522
  "0.95": {
10523
- "n": 471,
10524
- "errors": 38,
10525
- "error_rate": 0.08067940552016985
10526
  },
10527
  "0.99": {
10528
- "n": 303,
10529
- "errors": 10,
10530
- "error_rate": 0.033003300330033
10531
  }
10532
  },
10533
  "selective": {
10534
  "0.5": {
10535
  "coverage": 0.5,
10536
- "accuracy": 0.9024856596558317,
10537
- "confidence_cutoff": 0.9166792767745149
10538
  },
10539
  "0.8": {
10540
  "coverage": 0.8001912045889101,
10541
- "accuracy": 0.7765830346475507,
10542
- "confidence_cutoff": 0.5041101725134305
10543
  }
10544
  },
10545
  "score_mae": 0.35059676214309027,
@@ -10547,46 +10547,46 @@
10547
  },
10548
  "covered_only_probability_metrics": {
10549
  "n": 1046,
10550
- "nll": 0.9375416500577584,
10551
- "acc": 0.6797323135755258,
10552
- "ece": 0.10273742120233656,
10553
- "brier": 0.4350401558940114,
10554
- "mean_conf": 0.7714702411023553,
10555
- "confident_error_rate": 0.05449330783938815,
10556
- "coverage_at_0_9": 0.5239005736137667,
10557
- "accuracy_at_0_9": 0.8959854014598541,
10558
- "coverage_at_5pct_error": 0.35181644359464626,
10559
- "coverage_at_1pct_error": 0.14149139579349904,
10560
- "aurc": 0.11948583282589058,
10561
- "error_rate_at_0_9": 0.10401459854014598,
10562
- "confidence_bias": 0.09173792752682952,
10563
  "top_bins": {
10564
  "0.9": {
10565
- "n": 548,
10566
- "errors": 57,
10567
- "error_rate": 0.10401459854014598
10568
  },
10569
  "0.95": {
10570
- "n": 471,
10571
- "errors": 38,
10572
- "error_rate": 0.08067940552016985
10573
  },
10574
  "0.99": {
10575
- "n": 303,
10576
- "errors": 10,
10577
- "error_rate": 0.033003300330033
10578
  }
10579
  },
10580
  "selective": {
10581
  "0.5": {
10582
  "coverage": 0.5,
10583
- "accuracy": 0.9024856596558317,
10584
- "confidence_cutoff": 0.9166792767745149
10585
  },
10586
  "0.8": {
10587
  "coverage": 0.8001912045889101,
10588
- "accuracy": 0.7765830346475507,
10589
- "confidence_cutoff": 0.5041101725134305
10590
  }
10591
  },
10592
  "score_mae": 0.35059676214309027,
@@ -10598,33 +10598,33 @@
10598
  "buried_emotion": {
10599
  "requested": 20,
10600
  "valid_probability_rows": 20,
10601
- "correct": 4,
10602
- "accuracy": 0.2,
10603
  "full_probability_metrics": {
10604
  "n": 20,
10605
- "nll": 1.6915727437954797,
10606
- "acc": 0.2,
10607
- "ece": 0.09938172959866123,
10608
- "brier": 0.7953402434802076,
10609
- "mean_conf": 0.22966574304034268,
10610
  "confident_error_rate": 0.0,
10611
- "coverage_at_0_9": 0.0,
10612
- "accuracy_at_0_9": null,
10613
- "coverage_at_5pct_error": 0.1,
10614
- "coverage_at_1pct_error": 0.1,
10615
- "aurc": 0.611088792957988,
10616
- "error_rate_at_0_9": null,
10617
- "confidence_bias": 0.029665743040342668,
10618
  "top_bins": {
10619
  "0.9": {
10620
- "n": 0,
10621
  "errors": 0,
10622
- "error_rate": null
10623
  },
10624
  "0.95": {
10625
- "n": 0,
10626
  "errors": 0,
10627
- "error_rate": null
10628
  },
10629
  "0.99": {
10630
  "n": 0,
@@ -10635,41 +10635,41 @@
10635
  "selective": {
10636
  "0.5": {
10637
  "coverage": 0.5,
10638
- "accuracy": 0.3,
10639
- "confidence_cutoff": 0.23240028946100588
10640
  },
10641
  "0.8": {
10642
  "coverage": 0.8,
10643
- "accuracy": 0.25,
10644
- "confidence_cutoff": 0.19576863944530487
10645
  }
10646
  }
10647
  },
10648
  "covered_only_probability_metrics": {
10649
  "n": 20,
10650
- "nll": 1.6915727437954797,
10651
- "acc": 0.2,
10652
- "ece": 0.09938172959866123,
10653
- "brier": 0.7953402434802076,
10654
- "mean_conf": 0.22966574304034268,
10655
  "confident_error_rate": 0.0,
10656
- "coverage_at_0_9": 0.0,
10657
- "accuracy_at_0_9": null,
10658
- "coverage_at_5pct_error": 0.1,
10659
- "coverage_at_1pct_error": 0.1,
10660
- "aurc": 0.611088792957988,
10661
- "error_rate_at_0_9": null,
10662
- "confidence_bias": 0.029665743040342668,
10663
  "top_bins": {
10664
  "0.9": {
10665
- "n": 0,
10666
  "errors": 0,
10667
- "error_rate": null
10668
  },
10669
  "0.95": {
10670
- "n": 0,
10671
  "errors": 0,
10672
- "error_rate": null
10673
  },
10674
  "0.99": {
10675
  "n": 0,
@@ -10680,13 +10680,13 @@
10680
  "selective": {
10681
  "0.5": {
10682
  "coverage": 0.5,
10683
- "accuracy": 0.3,
10684
- "confidence_cutoff": 0.23240028946100588
10685
  },
10686
  "0.8": {
10687
  "coverage": 0.8,
10688
- "accuracy": 0.25,
10689
- "confidence_cutoff": 0.19576863944530487
10690
  }
10691
  }
10692
  },
@@ -11475,95 +11475,95 @@
11475
  "emotion": {
11476
  "requested": 80,
11477
  "valid_probability_rows": 80,
11478
- "correct": 32,
11479
- "accuracy": 0.4,
11480
  "full_probability_metrics": {
11481
  "n": 80,
11482
- "nll": 1.652865735056175,
11483
- "acc": 0.4,
11484
- "ece": 0.1426948414765427,
11485
- "brier": 0.7810760834350244,
11486
- "mean_conf": 0.2911186652438651,
11487
- "confident_error_rate": 0.0,
11488
- "coverage_at_0_9": 0.0,
11489
- "accuracy_at_0_9": null,
11490
- "coverage_at_5pct_error": 0.0,
11491
- "coverage_at_1pct_error": 0.0,
11492
- "aurc": 0.533927612449989,
11493
- "error_rate_at_0_9": null,
11494
- "confidence_bias": -0.10888133475613493,
11495
  "top_bins": {
11496
  "0.9": {
11497
- "n": 0,
11498
- "errors": 0,
11499
- "error_rate": null
11500
  },
11501
  "0.95": {
11502
- "n": 0,
11503
- "errors": 0,
11504
- "error_rate": null
11505
  },
11506
  "0.99": {
11507
- "n": 0,
11508
- "errors": 0,
11509
- "error_rate": null
11510
  }
11511
  },
11512
  "selective": {
11513
  "0.5": {
11514
  "coverage": 0.5,
11515
- "accuracy": 0.525,
11516
- "confidence_cutoff": 0.2595044115111889
11517
  },
11518
  "0.8": {
11519
  "coverage": 0.8,
11520
- "accuracy": 0.453125,
11521
- "confidence_cutoff": 0.21817853532039533
11522
  }
11523
  }
11524
  },
11525
  "covered_only_probability_metrics": {
11526
  "n": 80,
11527
- "nll": 1.652865735056175,
11528
- "acc": 0.4,
11529
- "ece": 0.1426948414765427,
11530
- "brier": 0.7810760834350244,
11531
- "mean_conf": 0.2911186652438651,
11532
- "confident_error_rate": 0.0,
11533
- "coverage_at_0_9": 0.0,
11534
- "accuracy_at_0_9": null,
11535
- "coverage_at_5pct_error": 0.0,
11536
- "coverage_at_1pct_error": 0.0,
11537
- "aurc": 0.533927612449989,
11538
- "error_rate_at_0_9": null,
11539
- "confidence_bias": -0.10888133475613493,
11540
  "top_bins": {
11541
  "0.9": {
11542
- "n": 0,
11543
- "errors": 0,
11544
- "error_rate": null
11545
  },
11546
  "0.95": {
11547
- "n": 0,
11548
- "errors": 0,
11549
- "error_rate": null
11550
  },
11551
  "0.99": {
11552
- "n": 0,
11553
- "errors": 0,
11554
- "error_rate": null
11555
  }
11556
  },
11557
  "selective": {
11558
  "0.5": {
11559
  "coverage": 0.5,
11560
- "accuracy": 0.525,
11561
- "confidence_cutoff": 0.2595044115111889
11562
  },
11563
  "0.8": {
11564
  "coverage": 0.8,
11565
- "accuracy": 0.453125,
11566
- "confidence_cutoff": 0.21817853532039533
11567
  }
11568
  }
11569
  },
@@ -13243,14 +13243,14 @@
13243
  "coverage": 1.0
13244
  }
13245
  },
13246
- "clean_task_macro_accuracy": 0.7204166666666667,
13247
  "permutation": {
13248
  "requested_pairs": 36,
13249
  "valid_pairs": 36,
13250
- "both_correct_requested": 0.5833333333333334,
13251
- "flip_rate": 0.25,
13252
- "covered_only_flip_rate": 0.25,
13253
- "half_l1": 0.10492792780864396
13254
  },
13255
  "unknowable": {
13256
  "requested": 110,
@@ -13378,7 +13378,7 @@
13378
  "truncated_questions": 0
13379
  },
13380
  "complete_unchanged_upstream_report": {
13381
- "objective": -0.8044143097078044,
13382
  "paired_flip": {
13383
  "pairs": 64,
13384
  "flip_rate": 0.828125,
@@ -13399,46 +13399,46 @@
13399
  },
13400
  "clean": {
13401
  "n": 1046,
13402
- "nll": 0.9375416500577584,
13403
- "acc": 0.6797323135755258,
13404
- "ece": 0.10273742120233656,
13405
- "brier": 0.4350401558940114,
13406
- "mean_conf": 0.7714702411023553,
13407
- "confident_error_rate": 0.05449330783938815,
13408
- "coverage_at_0_9": 0.5239005736137667,
13409
- "accuracy_at_0_9": 0.8959854014598541,
13410
- "coverage_at_5pct_error": 0.35181644359464626,
13411
- "coverage_at_1pct_error": 0.14149139579349904,
13412
- "aurc": 0.11948583282589058,
13413
- "error_rate_at_0_9": 0.10401459854014598,
13414
- "confidence_bias": 0.09173792752682952,
13415
  "top_bins": {
13416
  "0.9": {
13417
- "n": 548,
13418
- "errors": 57,
13419
- "error_rate": 0.10401459854014598
13420
  },
13421
  "0.95": {
13422
- "n": 471,
13423
- "errors": 38,
13424
- "error_rate": 0.08067940552016985
13425
  },
13426
  "0.99": {
13427
- "n": 303,
13428
- "errors": 10,
13429
- "error_rate": 0.033003300330033
13430
  }
13431
  },
13432
  "selective": {
13433
  "0.5": {
13434
  "coverage": 0.5,
13435
- "accuracy": 0.9024856596558317,
13436
- "confidence_cutoff": 0.9166792767745149
13437
  },
13438
  "0.8": {
13439
  "coverage": 0.8001912045889101,
13440
- "accuracy": 0.7765830346475507,
13441
- "confidence_cutoff": 0.5041101725134305
13442
  }
13443
  },
13444
  "score_mae": 0.35059676214309027,
@@ -13447,29 +13447,29 @@
13447
  "tasks": {
13448
  "buried_emotion": {
13449
  "n": 20,
13450
- "nll": 1.6915727437954797,
13451
- "acc": 0.2,
13452
- "ece": 0.09938172959866123,
13453
- "brier": 0.7953402434802076,
13454
- "mean_conf": 0.22966574304034268,
13455
  "confident_error_rate": 0.0,
13456
- "coverage_at_0_9": 0.0,
13457
- "accuracy_at_0_9": null,
13458
- "coverage_at_5pct_error": 0.1,
13459
- "coverage_at_1pct_error": 0.1,
13460
- "aurc": 0.611088792957988,
13461
- "error_rate_at_0_9": null,
13462
- "confidence_bias": 0.029665743040342668,
13463
  "top_bins": {
13464
  "0.9": {
13465
- "n": 0,
13466
  "errors": 0,
13467
- "error_rate": null
13468
  },
13469
  "0.95": {
13470
- "n": 0,
13471
  "errors": 0,
13472
- "error_rate": null
13473
  },
13474
  "0.99": {
13475
  "n": 0,
@@ -13480,13 +13480,13 @@
13480
  "selective": {
13481
  "0.5": {
13482
  "coverage": 0.5,
13483
- "accuracy": 0.3,
13484
- "confidence_cutoff": 0.23240028946100588
13485
  },
13486
  "0.8": {
13487
  "coverage": 0.8,
13488
- "accuracy": 0.25,
13489
- "confidence_cutoff": 0.19576863944530487
13490
  }
13491
  }
13492
  },
@@ -13854,46 +13854,46 @@
13854
  },
13855
  "emotion": {
13856
  "n": 80,
13857
- "nll": 1.652865735056175,
13858
- "acc": 0.4,
13859
- "ece": 0.1426948414765427,
13860
- "brier": 0.7810760834350244,
13861
- "mean_conf": 0.2911186652438651,
13862
- "confident_error_rate": 0.0,
13863
- "coverage_at_0_9": 0.0,
13864
- "accuracy_at_0_9": null,
13865
- "coverage_at_5pct_error": 0.0,
13866
- "coverage_at_1pct_error": 0.0,
13867
- "aurc": 0.533927612449989,
13868
- "error_rate_at_0_9": null,
13869
- "confidence_bias": -0.10888133475613493,
13870
  "top_bins": {
13871
  "0.9": {
13872
- "n": 0,
13873
- "errors": 0,
13874
- "error_rate": null
13875
  },
13876
  "0.95": {
13877
- "n": 0,
13878
- "errors": 0,
13879
- "error_rate": null
13880
  },
13881
  "0.99": {
13882
- "n": 0,
13883
- "errors": 0,
13884
- "error_rate": null
13885
  }
13886
  },
13887
  "selective": {
13888
  "0.5": {
13889
  "coverage": 0.5,
13890
- "accuracy": 0.525,
13891
- "confidence_cutoff": 0.2595044115111889
13892
  },
13893
  "0.8": {
13894
  "coverage": 0.8,
13895
- "accuracy": 0.453125,
13896
- "confidence_cutoff": 0.21817853532039533
13897
  }
13898
  }
13899
  },
@@ -15185,46 +15185,46 @@
15185
  "variants": {
15186
  "clean": {
15187
  "n": 1156,
15188
- "nll": 1.0275747126295693,
15189
- "acc": 0.6557093425605537,
15190
- "ece": 0.1271424265612776,
15191
- "brier": 0.4757092441503964,
15192
- "mean_conf": 0.7728989400694259,
15193
- "confident_error_rate": 0.06487889273356401,
15194
- "coverage_at_0_9": 0.5,
15195
- "accuracy_at_0_9": 0.870242214532872,
15196
- "coverage_at_5pct_error": 0.19896193771626297,
15197
- "coverage_at_1pct_error": 0.12802768166089964,
15198
- "aurc": 0.14592875492250384,
15199
- "error_rate_at_0_9": 0.12975778546712802,
15200
- "confidence_bias": 0.11718959750887226,
15201
- "top_bins": {
15202
- "0.9": {
15203
- "n": 578,
15204
- "errors": 75,
15205
- "error_rate": 0.12975778546712802
15206
- },
15207
- "0.95": {
15208
- "n": 495,
15209
- "errors": 51,
15210
- "error_rate": 0.10303030303030303
15211
  },
15212
  "0.99": {
15213
- "n": 317,
15214
- "errors": 18,
15215
- "error_rate": 0.056782334384858045
15216
  }
15217
  },
15218
  "selective": {
15219
  "0.5": {
15220
  "coverage": 0.5,
15221
- "accuracy": 0.870242214532872,
15222
- "confidence_cutoff": 0.9011673130850403
15223
  },
15224
  "0.8": {
15225
- "coverage": 0.8001730103806228,
15226
- "accuracy": 0.7437837837837837,
15227
- "confidence_cutoff": 0.5239011055369641
15228
  }
15229
  },
15230
  "score_mae": 0.49813993786469746,
@@ -15232,19 +15232,19 @@
15232
  },
15233
  "none_absent": {
15234
  "n": 36,
15235
- "nll": 1.9862789221192838,
15236
  "acc": 0.2777777777777778,
15237
- "ece": 0.36438230958721407,
15238
- "brier": 0.971223144388103,
15239
- "mean_conf": 0.5837732283166512,
15240
  "confident_error_rate": 0.05555555555555555,
15241
  "coverage_at_0_9": 0.1111111111111111,
15242
  "accuracy_at_0_9": 0.5,
15243
  "coverage_at_5pct_error": 0.0,
15244
  "coverage_at_1pct_error": 0.0,
15245
- "aurc": 0.6812529549976147,
15246
  "error_rate_at_0_9": 0.5,
15247
- "confidence_bias": 0.30599545053887345,
15248
  "top_bins": {
15249
  "0.9": {
15250
  "n": 4,
@@ -15265,39 +15265,39 @@
15265
  "selective": {
15266
  "0.5": {
15267
  "coverage": 0.5,
15268
- "accuracy": 0.3333333333333333,
15269
- "confidence_cutoff": 0.6306419821772263
15270
  },
15271
  "0.8": {
15272
  "coverage": 0.8055555555555556,
15273
- "accuracy": 0.27586206896551724,
15274
- "confidence_cutoff": 0.2793320479839753
15275
  }
15276
  }
15277
  },
15278
  "none_present": {
15279
  "n": 36,
15280
- "nll": 1.0208562024633812,
15281
- "acc": 0.6388888888888888,
15282
- "ece": 0.1927190351137436,
15283
- "brier": 0.5059127618759578,
15284
- "mean_conf": 0.6201323278766935,
15285
  "confident_error_rate": 0.0,
15286
- "coverage_at_0_9": 0.3333333333333333,
15287
  "accuracy_at_0_9": 1.0,
15288
- "coverage_at_5pct_error": 0.3888888888888889,
15289
- "coverage_at_1pct_error": 0.3888888888888889,
15290
- "aurc": 0.17333125994032075,
15291
  "error_rate_at_0_9": 0.0,
15292
- "confidence_bias": -0.01875656101219536,
15293
  "top_bins": {
15294
  "0.9": {
15295
- "n": 12,
15296
  "errors": 0,
15297
  "error_rate": 0.0
15298
  },
15299
  "0.95": {
15300
- "n": 11,
15301
  "errors": 0,
15302
  "error_rate": 0.0
15303
  },
@@ -15310,39 +15310,39 @@
15310
  "selective": {
15311
  "0.5": {
15312
  "coverage": 0.5,
15313
- "accuracy": 0.8333333333333334,
15314
- "confidence_cutoff": 0.6137714790540333
15315
  },
15316
  "0.8": {
15317
  "coverage": 0.8055555555555556,
15318
- "accuracy": 0.6551724137931034,
15319
- "confidence_cutoff": 0.30252245181248505
15320
  }
15321
  }
15322
  },
15323
  "permuted": {
15324
  "n": 36,
15325
- "nll": 0.9278919682020358,
15326
  "acc": 0.6666666666666666,
15327
- "ece": 0.16640228718149716,
15328
- "brier": 0.4734905986217854,
15329
- "mean_conf": 0.6290852939746607,
15330
  "confident_error_rate": 0.0,
15331
- "coverage_at_0_9": 0.3333333333333333,
15332
  "accuracy_at_0_9": 1.0,
15333
- "coverage_at_5pct_error": 0.3888888888888889,
15334
- "coverage_at_1pct_error": 0.3888888888888889,
15335
- "aurc": 0.15567120605578305,
15336
  "error_rate_at_0_9": 0.0,
15337
- "confidence_bias": -0.0375813726920059,
15338
  "top_bins": {
15339
  "0.9": {
15340
- "n": 12,
15341
  "errors": 0,
15342
  "error_rate": 0.0
15343
  },
15344
  "0.95": {
15345
- "n": 11,
15346
  "errors": 0,
15347
  "error_rate": 0.0
15348
  },
@@ -15355,13 +15355,13 @@
15355
  "selective": {
15356
  "0.5": {
15357
  "coverage": 0.5,
15358
- "accuracy": 0.8333333333333334,
15359
- "confidence_cutoff": 0.5823386094593029
15360
  },
15361
  "0.8": {
15362
  "coverage": 0.8055555555555556,
15363
- "accuracy": 0.6896551724137931,
15364
- "confidence_cutoff": 0.3432533187230439
15365
  }
15366
  }
15367
  }
@@ -15369,52 +15369,52 @@
15369
  "heldout_tasks": {},
15370
  "permutation": {
15371
  "n": 36,
15372
- "mean_max_delta": 0.08621858211804705,
15373
- "flip_rate": 0.25
15374
  },
15375
  "temperature": 1.0,
15376
  "calibrated_clean": {
15377
  "n": 1046,
15378
- "nll": 0.9375416500577584,
15379
- "acc": 0.6797323135755258,
15380
- "ece": 0.10273742120233656,
15381
- "brier": 0.4350401558940114,
15382
- "mean_conf": 0.7714702411023553,
15383
- "confident_error_rate": 0.05449330783938815,
15384
- "coverage_at_0_9": 0.5239005736137667,
15385
- "accuracy_at_0_9": 0.8959854014598541,
15386
- "coverage_at_5pct_error": 0.35181644359464626,
15387
- "coverage_at_1pct_error": 0.14149139579349904,
15388
- "aurc": 0.11948583282589058,
15389
- "error_rate_at_0_9": 0.10401459854014598,
15390
- "confidence_bias": 0.09173792752682952,
15391
  "top_bins": {
15392
  "0.9": {
15393
- "n": 548,
15394
- "errors": 57,
15395
- "error_rate": 0.10401459854014598
15396
  },
15397
  "0.95": {
15398
- "n": 471,
15399
- "errors": 38,
15400
- "error_rate": 0.08067940552016985
15401
  },
15402
  "0.99": {
15403
- "n": 303,
15404
- "errors": 10,
15405
- "error_rate": 0.033003300330033
15406
  }
15407
  },
15408
  "selective": {
15409
  "0.5": {
15410
  "coverage": 0.5,
15411
- "accuracy": 0.9024856596558317,
15412
- "confidence_cutoff": 0.9166792767745149
15413
  },
15414
  "0.8": {
15415
  "coverage": 0.8001912045889101,
15416
- "accuracy": 0.7765830346475507,
15417
- "confidence_cutoff": 0.5041101725134305
15418
  }
15419
  },
15420
  "score_mae": 0.35059676214309027,
@@ -66859,760 +66859,6 @@
66859
  }
66860
  },
66861
  "paired_comparisons": {
66862
- "Nox minus Sol": {
66863
- "old_core": {
66864
- "delta": 0.09255952380952381,
66865
- "exact_fraction": "311/3360",
66866
- "paired_ci95": [
66867
- 0.055577566964285605,
66868
- 0.1305831473214286
66869
- ]
66870
- },
66871
- "v3_core": {
66872
- "delta": 0.05708333333333333,
66873
- "exact_fraction": "137/2400",
66874
- "paired_ci95": [
66875
- 0.0245833333333334,
66876
- 0.08875
66877
- ]
66878
- },
66879
- "v4": {
66880
- "delta": 0.025,
66881
- "exact_fraction": "1/40",
66882
- "paired_ci95": [
66883
- -0.006288598445143156,
66884
- 0.058030063291139244
66885
- ]
66886
- },
66887
- "v5": {
66888
- "delta": 0.020833333333333332,
66889
- "exact_fraction": "1/48",
66890
- "paired_ci95": [
66891
- -0.006398928762210043,
66892
- 0.04862658709958234
66893
- ]
66894
- },
66895
- "transfer_v9_test": {
66896
- "delta": 0.1089866156787763,
66897
- "exact_fraction": "57/523",
66898
- "paired_ci95": [
66899
- 0.07972997762299643,
66900
- 0.13779904306220103
66901
- ]
66902
- },
66903
- "original_four_panel_mean": {
66904
- "delta": 0.04886904761904762,
66905
- "exact_fraction": "821/16800",
66906
- "paired_ci95": [
66907
- 0.03301257147338517,
66908
- 0.06502318612678312
66909
- ]
66910
- },
66911
- "weighted_mean": {
66912
- "delta": 0.06526168282800691,
66913
- "exact_fraction": "2293661/35145600",
66914
- "paired_ci95": [
66915
- 0.04958814518629331,
66916
- 0.08098195894796445
66917
- ]
66918
- }
66919
- },
66920
- "Nox minus kev-0.8b": {
66921
- "old_core": {
66922
- "delta": 0.22864583333333333,
66923
- "exact_fraction": "439/1920",
66924
- "paired_ci95": [
66925
- 0.17823567708333335,
66926
- 0.2778273809523808
66927
- ]
66928
- },
66929
- "v3_core": {
66930
- "delta": 0.095,
66931
- "exact_fraction": "19/200",
66932
- "paired_ci95": [
66933
- 0.0633333333333333,
66934
- 0.12791666666666673
66935
- ]
66936
- },
66937
- "v4": {
66938
- "delta": 0.1125,
66939
- "exact_fraction": "9/80",
66940
- "paired_ci95": [
66941
- 0.06741352201257866,
66942
- 0.15727848715504358
66943
- ]
66944
- },
66945
- "v5": {
66946
- "delta": 0.175,
66947
- "exact_fraction": "7/40",
66948
- "paired_ci95": [
66949
- 0.1355393511083524,
66950
- 0.21456521035365878
66951
- ]
66952
- },
66953
- "transfer_v9_test": {
66954
- "delta": 0.06787762906309751,
66955
- "exact_fraction": "71/1046",
66956
- "paired_ci95": [
66957
- 0.03431758137001224,
66958
- 0.10142248780822852
66959
- ]
66960
- },
66961
- "original_four_panel_mean": {
66962
- "delta": 0.15278645833333335,
66963
- "exact_fraction": "5867/38400",
66964
- "paired_ci95": [
66965
- 0.13186326764793416,
66966
- 0.1736005686577057
66967
- ]
66968
- },
66969
- "weighted_mean": {
66970
- "delta": 0.1456503943594646,
66971
- "exact_fraction": "487521/3347200",
66972
- "paired_ci95": [
66973
- 0.12568410699393018,
66974
- 0.16551635549679958
66975
- ]
66976
- }
66977
- },
66978
- "Nox minus kev-4b": {
66979
- "old_core": {
66980
- "delta": 0.11097470238095238,
66981
- "exact_fraction": "2983/26880",
66982
- "paired_ci95": [
66983
- 0.068154761904762,
66984
- 0.1539462425595238
66985
- ]
66986
- },
66987
- "v3_core": {
66988
- "delta": 0.0325,
66989
- "exact_fraction": "13/400",
66990
- "paired_ci95": [
66991
- 0.0033333333333334103,
66992
- 0.06166666666666676
66993
- ]
66994
- },
66995
- "v4": {
66996
- "delta": -0.028125,
66997
- "exact_fraction": "-9/320",
66998
- "paired_ci95": [
66999
- -0.062443225931676956,
67000
- 0.00613059214504551
67001
- ]
67002
- },
67003
- "v5": {
67004
- "delta": 0.016666666666666666,
67005
- "exact_fraction": "1/60",
67006
- "paired_ci95": [
67007
- -0.014652826016486136,
67008
- 0.04855467675421749
67009
- ]
67010
- },
67011
- "transfer_v9_test": {
67012
- "delta": -0.08126195028680688,
67013
- "exact_fraction": "-85/1046",
67014
- "paired_ci95": [
67015
- -0.1091449740378478,
67016
- -0.05362143534334946
67017
- ]
67018
- },
67019
- "original_four_panel_mean": {
67020
- "delta": 0.03300409226190476,
67021
- "exact_fraction": "17743/537600",
67022
- "paired_ci95": [
67023
- 0.0159171480237136,
67024
- 0.05071287650329332
67025
- ]
67026
- },
67027
- "weighted_mean": {
67028
- "delta": 0.027509368171264682,
67029
- "exact_fraction": "1289111/46860800",
67030
- "paired_ci95": [
67031
- 0.010595587102736993,
67032
- 0.044668398388024416
67033
- ]
67034
- }
67035
- },
67036
- "Nox minus kev-9b": {
67037
- "old_core": {
67038
- "delta": 0.06253720238095238,
67039
- "exact_fraction": "1681/26880",
67040
- "paired_ci95": [
67041
- 0.024514508928571384,
67042
- 0.10111700148809517
67043
- ]
67044
- },
67045
- "v3_core": {
67046
- "delta": 0.06041666666666667,
67047
- "exact_fraction": "29/480",
67048
- "paired_ci95": [
67049
- 0.02416666666666667,
67050
- 0.09708333333333341
67051
- ]
67052
- },
67053
- "v4": {
67054
- "delta": -0.0765625,
67055
- "exact_fraction": "-49/640",
67056
- "paired_ci95": [
67057
- -0.11406273964723933,
67058
- -0.040840863146287654
67059
- ]
67060
- },
67061
- "v5": {
67062
- "delta": 0.027083333333333334,
67063
- "exact_fraction": "13/480",
67064
- "paired_ci95": [
67065
- -0.008081049575002395,
67066
- 0.06178192588510811
67067
- ]
67068
- },
67069
- "transfer_v9_test": {
67070
- "delta": -0.11281070745697896,
67071
- "exact_fraction": "-59/523",
67072
- "paired_ci95": [
67073
- -0.14365179025963015,
67074
- -0.08235239701801159
67075
- ]
67076
- },
67077
- "original_four_panel_mean": {
67078
- "delta": 0.018368675595238096,
67079
- "exact_fraction": "395/21504",
67080
- "paired_ci95": [
67081
- 0.00015929724577673504,
67082
- 0.03635734921356825
67083
- ]
67084
- },
67085
- "weighted_mean": {
67086
- "delta": 0.009521846262405535,
67087
- "exact_fraction": "334651/35145600",
67088
- "paired_ci95": [
67089
- -0.007595435407328911,
67090
- 0.026801027422374526
67091
- ]
67092
- }
67093
- },
67094
- "Nox minus Jev": {
67095
- "old_core": {
67096
- "delta": 0.0390625,
67097
- "exact_fraction": "5/128",
67098
- "paired_ci95": [
67099
- 0.008703497023809521,
67100
- 0.06897321428571423
67101
- ]
67102
- },
67103
- "v3_core": {
67104
- "delta": -0.14583333333333334,
67105
- "exact_fraction": "-7/48",
67106
- "paired_ci95": [
67107
- -0.18834375000000017,
67108
- -0.10373958333333338
67109
- ]
67110
- },
67111
- "v4": {
67112
- "delta": -0.1546875,
67113
- "exact_fraction": "-99/640",
67114
- "paired_ci95": [
67115
- -0.19683761503067482,
67116
- -0.11406250000000007
67117
- ]
67118
- },
67119
- "v5": {
67120
- "delta": -0.035416666666666666,
67121
- "exact_fraction": "-17/480",
67122
- "paired_ci95": [
67123
- -0.06684327822543287,
67124
- -0.0039144015898180265
67125
- ]
67126
- },
67127
- "transfer_v9_test": {
67128
- "delta": -0.1921606118546845,
67129
- "exact_fraction": "-201/1046",
67130
- "paired_ci95": [
67131
- -0.2229480709592455,
67132
- -0.16162570888468808
67133
- ]
67134
- },
67135
- "original_four_panel_mean": {
67136
- "delta": -0.07421875,
67137
- "exact_fraction": "-19/256",
67138
- "paired_ci95": [
67139
- -0.09309387400474764,
67140
- -0.05577344994758131
67141
- ]
67142
- },
67143
- "weighted_mean": {
67144
- "delta": -0.08207930011153601,
67145
- "exact_fraction": "-329683/4016640",
67146
- "paired_ci95": [
67147
- -0.09913724619038332,
67148
- -0.06522481396228826
67149
- ]
67150
- }
67151
- },
67152
- "Nox minus Laya-base": {
67153
- "old_core": {
67154
- "delta": 0.26458333333333334,
67155
- "exact_fraction": "127/480",
67156
- "paired_ci95": [
67157
- 0.21428478422619038,
67158
- 0.31428664434523806
67159
- ]
67160
- },
67161
- "v3_core": {
67162
- "delta": 0.16458333333333333,
67163
- "exact_fraction": "79/480",
67164
- "paired_ci95": [
67165
- 0.12958333333333344,
67166
- 0.19874999999999993
67167
- ]
67168
- },
67169
- "v4": {
67170
- "delta": 0.2765625,
67171
- "exact_fraction": "177/640",
67172
- "paired_ci95": [
67173
- 0.2216732154792797,
67174
- 0.3312538580246914
67175
- ]
67176
- },
67177
- "v5": {
67178
- "delta": 0.225,
67179
- "exact_fraction": "9/40",
67180
- "paired_ci95": [
67181
- 0.18326279152965536,
67182
- 0.26607286866359453
67183
- ]
67184
- },
67185
- "transfer_v9_test": {
67186
- "delta": 0.1491395793499044,
67187
- "exact_fraction": "78/523",
67188
- "paired_ci95": [
67189
- 0.1103515625,
67190
- 0.18714025122546793
67191
- ]
67192
- },
67193
- "original_four_panel_mean": {
67194
- "delta": 0.23268229166666668,
67195
- "exact_fraction": "1787/7680",
67196
- "paired_ci95": [
67197
- 0.20941479053264894,
67198
- 0.2554643867871834
67199
- ]
67200
- },
67201
- "weighted_mean": {
67202
- "delta": 0.218126145235819,
67203
- "exact_fraction": "4380671/20083200",
67204
- "paired_ci95": [
67205
- 0.19712453859128337,
67206
- 0.23895587623050776
67207
- ]
67208
- }
67209
- },
67210
- "Nox minus Laya-multilingual": {
67211
- "old_core": {
67212
- "delta": 0.35751488095238093,
67213
- "exact_fraction": "961/2688",
67214
- "paired_ci95": [
67215
- 0.3083305431547619,
67216
- 0.40416666666666656
67217
- ]
67218
- },
67219
- "v3_core": {
67220
- "delta": 0.12875,
67221
- "exact_fraction": "103/800",
67222
- "paired_ci95": [
67223
- 0.09583333333333333,
67224
- 0.16166666666666668
67225
- ]
67226
- },
67227
- "v4": {
67228
- "delta": 0.2828125,
67229
- "exact_fraction": "181/640",
67230
- "paired_ci95": [
67231
- 0.2321340794423215,
67232
- 0.33590875589622643
67233
- ]
67234
- },
67235
- "v5": {
67236
- "delta": 0.28958333333333336,
67237
- "exact_fraction": "139/480",
67238
- "paired_ci95": [
67239
- 0.24203304326814726,
67240
- 0.33659227150537624
67241
- ]
67242
- },
67243
- "transfer_v9_test": {
67244
- "delta": 0.2084130019120459,
67245
- "exact_fraction": "109/523",
67246
- "paired_ci95": [
67247
- 0.16923040152963667,
67248
- 0.2462693477059149
67249
- ]
67250
- },
67251
- "original_four_panel_mean": {
67252
- "delta": 0.26466517857142857,
67253
- "exact_fraction": "11857/44800",
67254
- "paired_ci95": [
67255
- 0.24161685752065346,
67256
- 0.2875022282879109
67257
- ]
67258
- },
67259
- "weighted_mean": {
67260
- "delta": 0.25656328957252117,
67261
- "exact_fraction": "12022761/46860800",
67262
- "paired_ci95": [
67263
- 0.23600300131880433,
67264
- 0.27696035962801896
67265
- ]
67266
- }
67267
- },
67268
- "Nox minus Qwen3.5-9B": {
67269
- "old_core": {
67270
- "delta": 0.09088541666666666,
67271
- "exact_fraction": "349/3840",
67272
- "paired_ci95": [
67273
- 0.05066964285714293,
67274
- 0.13192057291666676
67275
- ]
67276
- },
67277
- "v3_core": {
67278
- "delta": 0.07166666666666667,
67279
- "exact_fraction": "43/600",
67280
- "paired_ci95": [
67281
- 0.038333333333333275,
67282
- 0.10499999999999993
67283
- ]
67284
- },
67285
- "v4": {
67286
- "delta": -0.1078125,
67287
- "exact_fraction": "-69/640",
67288
- "paired_ci95": [
67289
- -0.15406680969159642,
67290
- -0.0620528650542495
67291
- ]
67292
- },
67293
- "v5": {
67294
- "delta": 0.06666666666666667,
67295
- "exact_fraction": "1/15",
67296
- "paired_ci95": [
67297
- 0.027479948547215406,
67298
- 0.10597745181290595
67299
- ]
67300
- },
67301
- "transfer_v9_test": {
67302
- "delta": -0.052581261950286805,
67303
- "exact_fraction": "-55/1046",
67304
- "paired_ci95": [
67305
- -0.08596353013081821,
67306
- -0.01888529584691864
67307
- ]
67308
- },
67309
- "original_four_panel_mean": {
67310
- "delta": 0.0303515625,
67311
- "exact_fraction": "777/25600",
67312
- "paired_ci95": [
67313
- 0.01038540473033447,
67314
- 0.050420589249771885
67315
- ]
67316
- },
67317
- "weighted_mean": {
67318
- "delta": 0.031123227374123645,
67319
- "exact_fraction": "312527/10041600",
67320
- "paired_ci95": [
67321
- 0.013295818959364762,
67322
- 0.04931011228775257
67323
- ]
67324
- }
67325
- },
67326
- "Nox minus Lux": {
67327
- "old_core": {
67328
- "delta": -0.0020833333333333333,
67329
- "exact_fraction": "-1/480",
67330
- "paired_ci95": [
67331
- -0.03712797619047614,
67332
- 0.03303664434523808
67333
- ]
67334
- },
67335
- "v3_core": {
67336
- "delta": -0.0008333333333333334,
67337
- "exact_fraction": "-1/1200",
67338
- "paired_ci95": [
67339
- -0.0462499999999999,
67340
- 0.043749999999999956
67341
- ]
67342
- },
67343
- "v4": {
67344
- "delta": -0.1125,
67345
- "exact_fraction": "-9/80",
67346
- "paired_ci95": [
67347
- -0.15226071147798748,
67348
- -0.07448441096087467
67349
- ]
67350
- },
67351
- "v5": {
67352
- "delta": -0.04583333333333333,
67353
- "exact_fraction": "-11/240",
67354
- "paired_ci95": [
67355
- -0.07416056922725012,
67356
- -0.018749569559228588
67357
- ]
67358
- },
67359
- "transfer_v9_test": {
67360
- "delta": -0.09464627151051626,
67361
- "exact_fraction": "-99/1046",
67362
- "paired_ci95": [
67363
- -0.12101514862804885,
67364
- -0.06889873160603051
67365
- ]
67366
- },
67367
- "original_four_panel_mean": {
67368
- "delta": -0.0403125,
67369
- "exact_fraction": "-129/3200",
67370
- "paired_ci95": [
67371
- -0.05905983411444227,
67372
- -0.02179353585296119
67373
- ]
67374
- },
67375
- "weighted_mean": {
67376
- "delta": -0.038780274059910774,
67377
- "exact_fraction": "-48677/1255200",
67378
- "paired_ci95": [
67379
- -0.05625030602618532,
67380
- -0.021731519294624788
67381
- ]
67382
- }
67383
- },
67384
- "Nox minus Decider": {
67385
- "old_core": {
67386
- "delta": 0.18988095238095237,
67387
- "exact_fraction": "319/1680",
67388
- "paired_ci95": [
67389
- 0.1443443080357141,
67390
- 0.23571428571428554
67391
- ]
67392
- },
67393
- "v3_core": {
67394
- "delta": 0.052083333333333336,
67395
- "exact_fraction": "5/96",
67396
- "paired_ci95": [
67397
- 0.013333333333333308,
67398
- 0.08999999999999997
67399
- ]
67400
- },
67401
- "v4": {
67402
- "delta": -0.1296875,
67403
- "exact_fraction": "-83/640",
67404
- "paired_ci95": [
67405
- -0.17179976851851855,
67406
- -0.08906249999999993
67407
- ]
67408
- },
67409
- "v5": {
67410
- "delta": 0.01875,
67411
- "exact_fraction": "3/160",
67412
- "paired_ci95": [
67413
- -0.012535333369549428,
67414
- 0.04969926075268807
67415
- ]
67416
- },
67417
- "transfer_v9_test": {
67418
- "delta": -0.01338432122370937,
67419
- "exact_fraction": "-7/523",
67420
- "paired_ci95": [
67421
- -0.0490208104831727,
67422
- 0.022202716340089176
67423
- ]
67424
- },
67425
- "original_four_panel_mean": {
67426
- "delta": 0.03275669642857143,
67427
- "exact_fraction": "587/17920",
67428
- "paired_ci95": [
67429
- 0.013155496386521998,
67430
- 0.05228677715137045
67431
- ]
67432
- },
67433
- "weighted_mean": {
67434
- "delta": 0.05133684586406264,
67435
- "exact_fraction": "7217057/140582400",
67436
- "paired_ci95": [
67437
- 0.032083278884638945,
67438
- 0.07037086432355429
67439
- ]
67440
- }
67441
- },
67442
- "Nox minus Qwen3.5-2B": {
67443
- "old_core": {
67444
- "delta": 0.25877976190476193,
67445
- "exact_fraction": "1739/6720",
67446
- "paired_ci95": [
67447
- 0.21566220238095246,
67448
- 0.30174851190476193
67449
- ]
67450
- },
67451
- "v3_core": {
67452
- "delta": 0.12791666666666668,
67453
- "exact_fraction": "307/2400",
67454
- "paired_ci95": [
67455
- 0.0979166666666666,
67456
- 0.15999999999999992
67457
- ]
67458
- },
67459
- "v4": {
67460
- "delta": 0.053125,
67461
- "exact_fraction": "17/320",
67462
- "paired_ci95": [
67463
- 0.0011595366002794232,
67464
- 0.10381534584912663
67465
- ]
67466
- },
67467
- "v5": {
67468
- "delta": 0.13958333333333334,
67469
- "exact_fraction": "67/480",
67470
- "paired_ci95": [
67471
- 0.09772693403789712,
67472
- 0.18099321705426352
67473
- ]
67474
- },
67475
- "transfer_v9_test": {
67476
- "delta": 0.11663479923518165,
67477
- "exact_fraction": "61/523",
67478
- "paired_ci95": [
67479
- 0.08030488916814263,
67480
- 0.15300807906502473
67481
- ]
67482
- },
67483
- "original_four_panel_mean": {
67484
- "delta": 0.14485119047619047,
67485
- "exact_fraction": "4867/33600",
67486
- "paired_ci95": [
67487
- 0.12338022237889658,
67488
- 0.1658467989420549
67489
- ]
67490
- },
67491
- "weighted_mean": {
67492
- "delta": 0.15601456512337247,
67493
- "exact_fraction": "10966451/70291200",
67494
- "paired_ci95": [
67495
- 0.13723175354986636,
67496
- 0.17471297097636712
67497
- ]
67498
- }
67499
- },
67500
- "Nox minus Qwen3.5-4B": {
67501
- "old_core": {
67502
- "delta": 0.13113839285714285,
67503
- "exact_fraction": "235/1792",
67504
- "paired_ci95": [
67505
- 0.09270647321428567,
67506
- 0.16990420386904764
67507
- ]
67508
- },
67509
- "v3_core": {
67510
- "delta": 0.08458333333333333,
67511
- "exact_fraction": "203/2400",
67512
- "paired_ci95": [
67513
- 0.05416666666666664,
67514
- 0.11541666666666661
67515
- ]
67516
- },
67517
- "v4": {
67518
- "delta": -0.0890625,
67519
- "exact_fraction": "-57/640",
67520
- "paired_ci95": [
67521
- -0.1342959349117049,
67522
- -0.045753734276729595
67523
- ]
67524
- },
67525
- "v5": {
67526
- "delta": 0.06458333333333334,
67527
- "exact_fraction": "31/480",
67528
- "paired_ci95": [
67529
- 0.028211805555555566,
67530
- 0.10191039756413821
67531
- ]
67532
- },
67533
- "transfer_v9_test": {
67534
- "delta": -0.008604206500956023,
67535
- "exact_fraction": "-9/1046",
67536
- "paired_ci95": [
67537
- -0.040737677035593334,
67538
- 0.022878932316491855
67539
- ]
67540
- },
67541
- "original_four_panel_mean": {
67542
- "delta": 0.04781063988095238,
67543
- "exact_fraction": "25703/537600",
67544
- "paired_ci95": [
67545
- 0.028947119327501308,
67546
- 0.06683696978735856
67547
- ]
67548
- },
67549
- "weighted_mean": {
67550
- "delta": 0.05552484521533279,
67551
- "exact_fraction": "975727/17572800",
67552
- "paired_ci95": [
67553
- 0.03826384999615495,
67554
- 0.0725083343854517
67555
- ]
67556
- }
67557
- },
67558
- "Sol minus Nox": {
67559
- "old_core": {
67560
- "delta": -0.09255952380952381,
67561
- "exact_fraction": "-311/3360",
67562
- "paired_ci95": [
67563
- -0.13058314732142862,
67564
- -0.055577566964285646
67565
- ]
67566
- },
67567
- "v3_core": {
67568
- "delta": -0.05708333333333333,
67569
- "exact_fraction": "-137/2400",
67570
- "paired_ci95": [
67571
- -0.08875,
67572
- -0.0245833333333334
67573
- ]
67574
- },
67575
- "v4": {
67576
- "delta": -0.025,
67577
- "exact_fraction": "-1/40",
67578
- "paired_ci95": [
67579
- -0.058030063291139244,
67580
- 0.006288598445143156
67581
- ]
67582
- },
67583
- "v5": {
67584
- "delta": -0.020833333333333332,
67585
- "exact_fraction": "-1/48",
67586
- "paired_ci95": [
67587
- -0.048626587099582376,
67588
- 0.0063989287622100415
67589
- ]
67590
- },
67591
- "transfer_v9_test": {
67592
- "delta": -0.1089866156787763,
67593
- "exact_fraction": "-57/523",
67594
- "paired_ci95": [
67595
- -0.13779904306220103,
67596
- -0.07972997762299644
67597
- ]
67598
- },
67599
- "original_four_panel_mean": {
67600
- "delta": -0.04886904761904762,
67601
- "exact_fraction": "-821/16800",
67602
- "paired_ci95": [
67603
- -0.06502318612678314,
67604
- -0.033012571473385176
67605
- ]
67606
- },
67607
- "weighted_mean": {
67608
- "delta": -0.06526168282800691,
67609
- "exact_fraction": "-2293661/35145600",
67610
- "paired_ci95": [
67611
- -0.08098195894796448,
67612
- -0.049588145186293314
67613
- ]
67614
- }
67615
- },
67616
  "Sol minus kev-0.8b": {
67617
  "old_core": {
67618
  "delta": 0.1360863095238095,
@@ -68251,64 +67497,6 @@
68251
  ]
68252
  }
68253
  },
68254
- "Lux minus Nox": {
68255
- "old_core": {
68256
- "delta": 0.0020833333333333333,
68257
- "exact_fraction": "1/480",
68258
- "paired_ci95": [
68259
- -0.033036644345238085,
68260
- 0.03712797619047614
68261
- ]
68262
- },
68263
- "v3_core": {
68264
- "delta": 0.0008333333333333334,
68265
- "exact_fraction": "1/1200",
68266
- "paired_ci95": [
68267
- -0.043749999999999956,
68268
- 0.0462499999999999
68269
- ]
68270
- },
68271
- "v4": {
68272
- "delta": 0.1125,
68273
- "exact_fraction": "9/80",
68274
- "paired_ci95": [
68275
- 0.07448441096087466,
68276
- 0.15226071147798748
68277
- ]
68278
- },
68279
- "v5": {
68280
- "delta": 0.04583333333333333,
68281
- "exact_fraction": "11/240",
68282
- "paired_ci95": [
68283
- 0.018749569559228584,
68284
- 0.07416056922725009
68285
- ]
68286
- },
68287
- "transfer_v9_test": {
68288
- "delta": 0.09464627151051626,
68289
- "exact_fraction": "99/1046",
68290
- "paired_ci95": [
68291
- 0.0688987316060305,
68292
- 0.12101514862804882
68293
- ]
68294
- },
68295
- "original_four_panel_mean": {
68296
- "delta": 0.0403125,
68297
- "exact_fraction": "129/3200",
68298
- "paired_ci95": [
68299
- 0.02179353585296119,
68300
- 0.05905983411444227
68301
- ]
68302
- },
68303
- "weighted_mean": {
68304
- "delta": 0.038780274059910774,
68305
- "exact_fraction": "48677/1255200",
68306
- "paired_ci95": [
68307
- 0.021731519294624784,
68308
- 0.05625030602618531
68309
- ]
68310
- }
68311
- },
68312
  "Lux minus Sol": {
68313
  "old_core": {
68314
  "delta": 0.09464285714285714,
@@ -68946,6 +68134,180 @@
68946
  0.11044819701768911
68947
  ]
68948
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
68949
  }
68950
  },
68951
  "method": {
@@ -68973,10 +68335,10 @@
68973
  "transfer_accuracy": "Upstream ordered probability argmax, clean knowable records, micro average. Refused/missing probabilities count wrong; unknown evidence is not accuracy.",
68974
  "transfer_clusters": "Union IDs/group/parent/control/pair plus exact state SHA and ordered-request IDs, built over all variants then scored only clean knowable rows. Complete components jointly resampled; variable sampled denominator. No whole-family grouping.",
68975
  "pairing": "Identical sampled source components across all models; five panels sampled independently. Original four marginal CIs preserved.",
68976
- "accuracy_arithmetic": "Exact rational integer correct/requested. Product weights were explicitly chosen after observing earlier outcomes; no claim of outcome-independent weighting.",
68977
  "interval_role": "Descriptive paired95% intervals; no positive-CI release requirement invented.",
68978
  "diagnostics": "Full per-task/proper-score/order/unknown-evidence results remain separate; no accuracy/latency/ECE unit mixing.",
68979
- "exposure": "Observed regression data; latest30/25/15/15/15 weighting is outcome-informed user steering. Identical predictions, no test-based checkpoint reselection or calibration."
68980
  },
68981
  "prior_weighted_scores": {
68982
  "Jev": {
@@ -68996,12 +68358,9 @@
68996
  ]
68997
  },
68998
  "Nox": {
68999
- "point": 0.7208999597389147,
69000
- "exact_fraction": "202691693/281164800",
69001
- "ci95": [
69002
- 0.706127064217637,
69003
- 0.7352055352744082
69004
- ]
69005
  },
69006
  "kev-9b": {
69007
  "point": 0.7201455089684057,
@@ -69086,6 +68445,6 @@
69086
  },
69087
  "hosted_frontier": "All official calls returned jev-1.13.0; closed weight identity/input retention cannot be independently inspected.",
69088
  "size_note": "Qwen3.5 family size labels name their official parent checkpoint. Lux deployed text+decision weights total7,940,895,744 parameters; no vision or generativeLMhead in its serving bundle.",
69089
- "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f",
69090
- "presentation_note": "Public display roster changed only; frozen protocol, complete internal evidence and every retained numeric result are unchanged."
69091
  }
 
9
  },
10
  "metric_design": "User-requested, outcome-informed decision-priority weights; observed regression data, not blind testing.",
11
  "protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
12
+ "statistics_sha256": "b319a993ac569ea1fa8da027ccb2195868557e4efe232a5d06091db2884086be",
13
  "models": {
14
  "Jev": {
15
  "panels": {
 
10330
  ]
10331
  },
10332
  "transfer_v9_test": {
10333
+ "point": 0.6959847036328872,
10334
+ "exact_fraction": "364/523",
10335
  "ci95": [
10336
+ 0.6660212501997199,
10337
+ 0.7249533151094627
10338
  ]
10339
  },
10340
  "original_four_panel_mean": {
 
10346
  ]
10347
  },
10348
  "weighted_mean": {
10349
+ "point": 0.7308523186401712,
10350
+ "exact_fraction": "102744973/140582400",
10351
  "ci95": [
10352
+ 0.7157461705500261,
10353
+ 0.7455873812383725
10354
  ]
10355
  }
10356
  },
 
10496
  "clean": {
10497
  "requested": 1046,
10498
  "valid_probability_rows": 1046,
10499
+ "correct": 728,
10500
+ "accuracy": 0.6959847036328872,
10501
  "full_probability_metrics": {
10502
  "n": 1046,
10503
+ "nll": 0.9079076925806586,
10504
+ "acc": 0.6959847036328872,
10505
+ "ece": 0.11803573910263573,
10506
+ "brier": 0.41689550886050714,
10507
+ "mean_conf": 0.814020442735523,
10508
+ "confident_error_rate": 0.058317399617590825,
10509
+ "coverage_at_0_9": 0.5525812619502868,
10510
+ "accuracy_at_0_9": 0.8944636678200693,
10511
+ "coverage_at_5pct_error": 0.36424474187380496,
10512
+ "coverage_at_1pct_error": 0.14340344168260039,
10513
+ "aurc": 0.11088769274123257,
10514
+ "error_rate_at_0_9": 0.10553633217993079,
10515
+ "confidence_bias": 0.11803573910263576,
10516
  "top_bins": {
10517
  "0.9": {
10518
+ "n": 578,
10519
+ "errors": 61,
10520
+ "error_rate": 0.10553633217993079
10521
  },
10522
  "0.95": {
10523
+ "n": 491,
10524
+ "errors": 39,
10525
+ "error_rate": 0.07942973523421588
10526
  },
10527
  "0.99": {
10528
+ "n": 314,
10529
+ "errors": 11,
10530
+ "error_rate": 0.03503184713375796
10531
  }
10532
  },
10533
  "selective": {
10534
  "0.5": {
10535
  "coverage": 0.5,
10536
+ "accuracy": 0.9101338432122371,
10537
+ "confidence_cutoff": 0.9301925551074537
10538
  },
10539
  "0.8": {
10540
  "coverage": 0.8001912045889101,
10541
+ "accuracy": 0.7921146953405018,
10542
+ "confidence_cutoff": 0.5957019462220806
10543
  }
10544
  },
10545
  "score_mae": 0.35059676214309027,
 
10547
  },
10548
  "covered_only_probability_metrics": {
10549
  "n": 1046,
10550
+ "nll": 0.9079076925806586,
10551
+ "acc": 0.6959847036328872,
10552
+ "ece": 0.11803573910263573,
10553
+ "brier": 0.41689550886050714,
10554
+ "mean_conf": 0.814020442735523,
10555
+ "confident_error_rate": 0.058317399617590825,
10556
+ "coverage_at_0_9": 0.5525812619502868,
10557
+ "accuracy_at_0_9": 0.8944636678200693,
10558
+ "coverage_at_5pct_error": 0.36424474187380496,
10559
+ "coverage_at_1pct_error": 0.14340344168260039,
10560
+ "aurc": 0.11088769274123257,
10561
+ "error_rate_at_0_9": 0.10553633217993079,
10562
+ "confidence_bias": 0.11803573910263576,
10563
  "top_bins": {
10564
  "0.9": {
10565
+ "n": 578,
10566
+ "errors": 61,
10567
+ "error_rate": 0.10553633217993079
10568
  },
10569
  "0.95": {
10570
+ "n": 491,
10571
+ "errors": 39,
10572
+ "error_rate": 0.07942973523421588
10573
  },
10574
  "0.99": {
10575
+ "n": 314,
10576
+ "errors": 11,
10577
+ "error_rate": 0.03503184713375796
10578
  }
10579
  },
10580
  "selective": {
10581
  "0.5": {
10582
  "coverage": 0.5,
10583
+ "accuracy": 0.9101338432122371,
10584
+ "confidence_cutoff": 0.9301925551074537
10585
  },
10586
  "0.8": {
10587
  "coverage": 0.8001912045889101,
10588
+ "accuracy": 0.7921146953405018,
10589
+ "confidence_cutoff": 0.5957019462220806
10590
  }
10591
  },
10592
  "score_mae": 0.35059676214309027,
 
10598
  "buried_emotion": {
10599
  "requested": 20,
10600
  "valid_probability_rows": 20,
10601
+ "correct": 5,
10602
+ "accuracy": 0.25,
10603
  "full_probability_metrics": {
10604
  "n": 20,
10605
+ "nll": 1.4541223366826688,
10606
+ "acc": 0.25,
10607
+ "ece": 0.33647159228575807,
10608
+ "brier": 0.784234954627584,
10609
+ "mean_conf": 0.5836932166148608,
10610
  "confident_error_rate": 0.0,
10611
+ "coverage_at_0_9": 0.05,
10612
+ "accuracy_at_0_9": 1.0,
10613
+ "coverage_at_5pct_error": 0.05,
10614
+ "coverage_at_1pct_error": 0.05,
10615
+ "aurc": 0.5585412761902699,
10616
+ "error_rate_at_0_9": 0.0,
10617
+ "confidence_bias": 0.33369321661486084,
10618
  "top_bins": {
10619
  "0.9": {
10620
+ "n": 1,
10621
  "errors": 0,
10622
+ "error_rate": 0.0
10623
  },
10624
  "0.95": {
10625
+ "n": 1,
10626
  "errors": 0,
10627
+ "error_rate": 0.0
10628
  },
10629
  "0.99": {
10630
  "n": 0,
 
10635
  "selective": {
10636
  "0.5": {
10637
  "coverage": 0.5,
10638
+ "accuracy": 0.5,
10639
+ "confidence_cutoff": 0.5472205811051023
10640
  },
10641
  "0.8": {
10642
  "coverage": 0.8,
10643
+ "accuracy": 0.3125,
10644
+ "confidence_cutoff": 0.4314323265237658
10645
  }
10646
  }
10647
  },
10648
  "covered_only_probability_metrics": {
10649
  "n": 20,
10650
+ "nll": 1.4541223366826688,
10651
+ "acc": 0.25,
10652
+ "ece": 0.33647159228575807,
10653
+ "brier": 0.784234954627584,
10654
+ "mean_conf": 0.5836932166148608,
10655
  "confident_error_rate": 0.0,
10656
+ "coverage_at_0_9": 0.05,
10657
+ "accuracy_at_0_9": 1.0,
10658
+ "coverage_at_5pct_error": 0.05,
10659
+ "coverage_at_1pct_error": 0.05,
10660
+ "aurc": 0.5585412761902699,
10661
+ "error_rate_at_0_9": 0.0,
10662
+ "confidence_bias": 0.33369321661486084,
10663
  "top_bins": {
10664
  "0.9": {
10665
+ "n": 1,
10666
  "errors": 0,
10667
+ "error_rate": 0.0
10668
  },
10669
  "0.95": {
10670
+ "n": 1,
10671
  "errors": 0,
10672
+ "error_rate": 0.0
10673
  },
10674
  "0.99": {
10675
  "n": 0,
 
10680
  "selective": {
10681
  "0.5": {
10682
  "coverage": 0.5,
10683
+ "accuracy": 0.5,
10684
+ "confidence_cutoff": 0.5472205811051023
10685
  },
10686
  "0.8": {
10687
  "coverage": 0.8,
10688
+ "accuracy": 0.3125,
10689
+ "confidence_cutoff": 0.4314323265237658
10690
  }
10691
  }
10692
  },
 
11475
  "emotion": {
11476
  "requested": 80,
11477
  "valid_probability_rows": 80,
11478
+ "correct": 48,
11479
+ "accuracy": 0.6,
11480
  "full_probability_metrics": {
11481
  "n": 80,
11482
+ "nll": 1.3247643428213003,
11483
+ "acc": 0.6,
11484
+ "ece": 0.15895568320390185,
11485
+ "brier": 0.5466111456851124,
11486
+ "mean_conf": 0.7589556832039019,
11487
+ "confident_error_rate": 0.05,
11488
+ "coverage_at_0_9": 0.3625,
11489
+ "accuracy_at_0_9": 0.8620689655172413,
11490
+ "coverage_at_5pct_error": 0.25,
11491
+ "coverage_at_1pct_error": 0.0375,
11492
+ "aurc": 0.20182699408158983,
11493
+ "error_rate_at_0_9": 0.13793103448275862,
11494
+ "confidence_bias": 0.15895568320390197,
11495
  "top_bins": {
11496
  "0.9": {
11497
+ "n": 29,
11498
+ "errors": 4,
11499
+ "error_rate": 0.13793103448275862
11500
  },
11501
  "0.95": {
11502
+ "n": 19,
11503
+ "errors": 1,
11504
+ "error_rate": 0.05263157894736842
11505
  },
11506
  "0.99": {
11507
+ "n": 11,
11508
+ "errors": 1,
11509
+ "error_rate": 0.09090909090909091
11510
  }
11511
  },
11512
  "selective": {
11513
  "0.5": {
11514
  "coverage": 0.5,
11515
+ "accuracy": 0.85,
11516
+ "confidence_cutoff": 0.7988864004860878
11517
  },
11518
  "0.8": {
11519
  "coverage": 0.8,
11520
+ "accuracy": 0.671875,
11521
+ "confidence_cutoff": 0.5632915364363126
11522
  }
11523
  }
11524
  },
11525
  "covered_only_probability_metrics": {
11526
  "n": 80,
11527
+ "nll": 1.3247643428213003,
11528
+ "acc": 0.6,
11529
+ "ece": 0.15895568320390185,
11530
+ "brier": 0.5466111456851124,
11531
+ "mean_conf": 0.7589556832039019,
11532
+ "confident_error_rate": 0.05,
11533
+ "coverage_at_0_9": 0.3625,
11534
+ "accuracy_at_0_9": 0.8620689655172413,
11535
+ "coverage_at_5pct_error": 0.25,
11536
+ "coverage_at_1pct_error": 0.0375,
11537
+ "aurc": 0.20182699408158983,
11538
+ "error_rate_at_0_9": 0.13793103448275862,
11539
+ "confidence_bias": 0.15895568320390197,
11540
  "top_bins": {
11541
  "0.9": {
11542
+ "n": 29,
11543
+ "errors": 4,
11544
+ "error_rate": 0.13793103448275862
11545
  },
11546
  "0.95": {
11547
+ "n": 19,
11548
+ "errors": 1,
11549
+ "error_rate": 0.05263157894736842
11550
  },
11551
  "0.99": {
11552
+ "n": 11,
11553
+ "errors": 1,
11554
+ "error_rate": 0.09090909090909091
11555
  }
11556
  },
11557
  "selective": {
11558
  "0.5": {
11559
  "coverage": 0.5,
11560
+ "accuracy": 0.85,
11561
+ "confidence_cutoff": 0.7988864004860878
11562
  },
11563
  "0.8": {
11564
  "coverage": 0.8,
11565
+ "accuracy": 0.671875,
11566
+ "confidence_cutoff": 0.5632915364363126
11567
  }
11568
  }
11569
  },
 
13243
  "coverage": 1.0
13244
  }
13245
  },
13246
+ "clean_task_macro_accuracy": 0.729675925925926,
13247
  "permutation": {
13248
  "requested_pairs": 36,
13249
  "valid_pairs": 36,
13250
+ "both_correct_requested": 0.6388888888888888,
13251
+ "flip_rate": 0.16666666666666666,
13252
+ "covered_only_flip_rate": 0.16666666666666666,
13253
+ "half_l1": 0.07949225370274392
13254
  },
13255
  "unknowable": {
13256
  "requested": 110,
 
13378
  "truncated_questions": 0
13379
  },
13380
  "complete_unchanged_upstream_report": {
13381
+ "objective": -0.7834679467690011,
13382
  "paired_flip": {
13383
  "pairs": 64,
13384
  "flip_rate": 0.828125,
 
13399
  },
13400
  "clean": {
13401
  "n": 1046,
13402
+ "nll": 0.9079076925806586,
13403
+ "acc": 0.6959847036328872,
13404
+ "ece": 0.11803573910263573,
13405
+ "brier": 0.41689550886050714,
13406
+ "mean_conf": 0.814020442735523,
13407
+ "confident_error_rate": 0.058317399617590825,
13408
+ "coverage_at_0_9": 0.5525812619502868,
13409
+ "accuracy_at_0_9": 0.8944636678200693,
13410
+ "coverage_at_5pct_error": 0.36424474187380496,
13411
+ "coverage_at_1pct_error": 0.14340344168260039,
13412
+ "aurc": 0.11088769274123257,
13413
+ "error_rate_at_0_9": 0.10553633217993079,
13414
+ "confidence_bias": 0.11803573910263576,
13415
  "top_bins": {
13416
  "0.9": {
13417
+ "n": 578,
13418
+ "errors": 61,
13419
+ "error_rate": 0.10553633217993079
13420
  },
13421
  "0.95": {
13422
+ "n": 491,
13423
+ "errors": 39,
13424
+ "error_rate": 0.07942973523421588
13425
  },
13426
  "0.99": {
13427
+ "n": 314,
13428
+ "errors": 11,
13429
+ "error_rate": 0.03503184713375796
13430
  }
13431
  },
13432
  "selective": {
13433
  "0.5": {
13434
  "coverage": 0.5,
13435
+ "accuracy": 0.9101338432122371,
13436
+ "confidence_cutoff": 0.9301925551074537
13437
  },
13438
  "0.8": {
13439
  "coverage": 0.8001912045889101,
13440
+ "accuracy": 0.7921146953405018,
13441
+ "confidence_cutoff": 0.5957019462220806
13442
  }
13443
  },
13444
  "score_mae": 0.35059676214309027,
 
13447
  "tasks": {
13448
  "buried_emotion": {
13449
  "n": 20,
13450
+ "nll": 1.4541223366826688,
13451
+ "acc": 0.25,
13452
+ "ece": 0.33647159228575807,
13453
+ "brier": 0.784234954627584,
13454
+ "mean_conf": 0.5836932166148608,
13455
  "confident_error_rate": 0.0,
13456
+ "coverage_at_0_9": 0.05,
13457
+ "accuracy_at_0_9": 1.0,
13458
+ "coverage_at_5pct_error": 0.05,
13459
+ "coverage_at_1pct_error": 0.05,
13460
+ "aurc": 0.5585412761902699,
13461
+ "error_rate_at_0_9": 0.0,
13462
+ "confidence_bias": 0.33369321661486084,
13463
  "top_bins": {
13464
  "0.9": {
13465
+ "n": 1,
13466
  "errors": 0,
13467
+ "error_rate": 0.0
13468
  },
13469
  "0.95": {
13470
+ "n": 1,
13471
  "errors": 0,
13472
+ "error_rate": 0.0
13473
  },
13474
  "0.99": {
13475
  "n": 0,
 
13480
  "selective": {
13481
  "0.5": {
13482
  "coverage": 0.5,
13483
+ "accuracy": 0.5,
13484
+ "confidence_cutoff": 0.5472205811051023
13485
  },
13486
  "0.8": {
13487
  "coverage": 0.8,
13488
+ "accuracy": 0.3125,
13489
+ "confidence_cutoff": 0.4314323265237658
13490
  }
13491
  }
13492
  },
 
13854
  },
13855
  "emotion": {
13856
  "n": 80,
13857
+ "nll": 1.3247643428213003,
13858
+ "acc": 0.6,
13859
+ "ece": 0.15895568320390185,
13860
+ "brier": 0.5466111456851124,
13861
+ "mean_conf": 0.7589556832039019,
13862
+ "confident_error_rate": 0.05,
13863
+ "coverage_at_0_9": 0.3625,
13864
+ "accuracy_at_0_9": 0.8620689655172413,
13865
+ "coverage_at_5pct_error": 0.25,
13866
+ "coverage_at_1pct_error": 0.0375,
13867
+ "aurc": 0.20182699408158983,
13868
+ "error_rate_at_0_9": 0.13793103448275862,
13869
+ "confidence_bias": 0.15895568320390197,
13870
  "top_bins": {
13871
  "0.9": {
13872
+ "n": 29,
13873
+ "errors": 4,
13874
+ "error_rate": 0.13793103448275862
13875
  },
13876
  "0.95": {
13877
+ "n": 19,
13878
+ "errors": 1,
13879
+ "error_rate": 0.05263157894736842
13880
  },
13881
  "0.99": {
13882
+ "n": 11,
13883
+ "errors": 1,
13884
+ "error_rate": 0.09090909090909091
13885
  }
13886
  },
13887
  "selective": {
13888
  "0.5": {
13889
  "coverage": 0.5,
13890
+ "accuracy": 0.85,
13891
+ "confidence_cutoff": 0.7988864004860878
13892
  },
13893
  "0.8": {
13894
  "coverage": 0.8,
13895
+ "accuracy": 0.671875,
13896
+ "confidence_cutoff": 0.5632915364363126
13897
  }
13898
  }
13899
  },
 
15185
  "variants": {
15186
  "clean": {
15187
  "n": 1156,
15188
+ "nll": 1.0007605953968302,
15189
+ "acc": 0.6704152249134948,
15190
+ "ece": 0.1409850221700258,
15191
+ "brier": 0.45929116387613567,
15192
+ "mean_conf": 0.8114002470835204,
15193
+ "confident_error_rate": 0.06833910034602077,
15194
+ "coverage_at_0_9": 0.5259515570934256,
15195
+ "accuracy_at_0_9": 0.8700657894736842,
15196
+ "coverage_at_5pct_error": 0.20242214532871972,
15197
+ "coverage_at_1pct_error": 0.12975778546712802,
15198
+ "aurc": 0.1369722577685153,
15199
+ "error_rate_at_0_9": 0.1299342105263158,
15200
+ "confidence_bias": 0.14098502217002562,
15201
+ "top_bins": {
15202
+ "0.9": {
15203
+ "n": 608,
15204
+ "errors": 79,
15205
+ "error_rate": 0.1299342105263158
15206
+ },
15207
+ "0.95": {
15208
+ "n": 515,
15209
+ "errors": 52,
15210
+ "error_rate": 0.10097087378640776
15211
  },
15212
  "0.99": {
15213
+ "n": 328,
15214
+ "errors": 19,
15215
+ "error_rate": 0.057926829268292686
15216
  }
15217
  },
15218
  "selective": {
15219
  "0.5": {
15220
  "coverage": 0.5,
15221
+ "accuracy": 0.8788927335640139,
15222
+ "confidence_cutoff": 0.9157804353063602
15223
  },
15224
  "0.8": {
15225
+ "coverage": 0.801038062283737,
15226
+ "accuracy": 0.7591792656587473,
15227
+ "confidence_cutoff": 0.6007340833161877
15228
  }
15229
  },
15230
  "score_mae": 0.49813993786469746,
 
15232
  },
15233
  "none_absent": {
15234
  "n": 36,
15235
+ "nll": 2.016479899399725,
15236
  "acc": 0.2777777777777778,
15237
+ "ece": 0.43639357499087156,
15238
+ "brier": 1.0193378435955023,
15239
+ "mean_conf": 0.7141713527686493,
15240
  "confident_error_rate": 0.05555555555555555,
15241
  "coverage_at_0_9": 0.1111111111111111,
15242
  "accuracy_at_0_9": 0.5,
15243
  "coverage_at_5pct_error": 0.0,
15244
  "coverage_at_1pct_error": 0.0,
15245
+ "aurc": 0.6622061759903008,
15246
  "error_rate_at_0_9": 0.5,
15247
+ "confidence_bias": 0.4363935749908715,
15248
  "top_bins": {
15249
  "0.9": {
15250
  "n": 4,
 
15265
  "selective": {
15266
  "0.5": {
15267
  "coverage": 0.5,
15268
+ "accuracy": 0.3888888888888889,
15269
+ "confidence_cutoff": 0.7284490118427092
15270
  },
15271
  "0.8": {
15272
  "coverage": 0.8055555555555556,
15273
+ "accuracy": 0.3448275862068966,
15274
+ "confidence_cutoff": 0.5772219386282852
15275
  }
15276
  }
15277
  },
15278
  "none_present": {
15279
  "n": 36,
15280
+ "nll": 0.9834989953062324,
15281
+ "acc": 0.6666666666666666,
15282
+ "ece": 0.16246789063665687,
15283
+ "brier": 0.42631332906431735,
15284
+ "mean_conf": 0.7475449294183114,
15285
  "confident_error_rate": 0.0,
15286
+ "coverage_at_0_9": 0.3888888888888889,
15287
  "accuracy_at_0_9": 1.0,
15288
+ "coverage_at_5pct_error": 0.5,
15289
+ "coverage_at_1pct_error": 0.5,
15290
+ "aurc": 0.10954859179448669,
15291
  "error_rate_at_0_9": 0.0,
15292
+ "confidence_bias": 0.08087826275164478,
15293
  "top_bins": {
15294
  "0.9": {
15295
+ "n": 14,
15296
  "errors": 0,
15297
  "error_rate": 0.0
15298
  },
15299
  "0.95": {
15300
+ "n": 12,
15301
  "errors": 0,
15302
  "error_rate": 0.0
15303
  },
 
15310
  "selective": {
15311
  "0.5": {
15312
  "coverage": 0.5,
15313
+ "accuracy": 1.0,
15314
+ "confidence_cutoff": 0.8038808995404794
15315
  },
15316
  "0.8": {
15317
  "coverage": 0.8055555555555556,
15318
+ "accuracy": 0.7586206896551724,
15319
+ "confidence_cutoff": 0.514763438695478
15320
  }
15321
  }
15322
  },
15323
  "permuted": {
15324
  "n": 36,
15325
+ "nll": 0.9187906169822195,
15326
  "acc": 0.6666666666666666,
15327
+ "ece": 0.13776171917514016,
15328
+ "brier": 0.4066078577529032,
15329
+ "mean_conf": 0.7543481885836296,
15330
  "confident_error_rate": 0.0,
15331
+ "coverage_at_0_9": 0.4166666666666667,
15332
  "accuracy_at_0_9": 1.0,
15333
+ "coverage_at_5pct_error": 0.5555555555555556,
15334
+ "coverage_at_1pct_error": 0.5,
15335
+ "aurc": 0.09826261699381092,
15336
  "error_rate_at_0_9": 0.0,
15337
+ "confidence_bias": 0.08768152191696299,
15338
  "top_bins": {
15339
  "0.9": {
15340
+ "n": 15,
15341
  "errors": 0,
15342
  "error_rate": 0.0
15343
  },
15344
  "0.95": {
15345
+ "n": 14,
15346
  "errors": 0,
15347
  "error_rate": 0.0
15348
  },
 
15355
  "selective": {
15356
  "0.5": {
15357
  "coverage": 0.5,
15358
+ "accuracy": 1.0,
15359
+ "confidence_cutoff": 0.8030261229089575
15360
  },
15361
  "0.8": {
15362
  "coverage": 0.8055555555555556,
15363
+ "accuracy": 0.7931034482758621,
15364
+ "confidence_cutoff": 0.5168578524128333
15365
  }
15366
  }
15367
  }
 
15369
  "heldout_tasks": {},
15370
  "permutation": {
15371
  "n": 36,
15372
+ "mean_max_delta": 0.07510801427468833,
15373
+ "flip_rate": 0.16666666666666666
15374
  },
15375
  "temperature": 1.0,
15376
  "calibrated_clean": {
15377
  "n": 1046,
15378
+ "nll": 0.9079076925806586,
15379
+ "acc": 0.6959847036328872,
15380
+ "ece": 0.11803573910263573,
15381
+ "brier": 0.41689550886050714,
15382
+ "mean_conf": 0.814020442735523,
15383
+ "confident_error_rate": 0.058317399617590825,
15384
+ "coverage_at_0_9": 0.5525812619502868,
15385
+ "accuracy_at_0_9": 0.8944636678200693,
15386
+ "coverage_at_5pct_error": 0.36424474187380496,
15387
+ "coverage_at_1pct_error": 0.14340344168260039,
15388
+ "aurc": 0.11088769274123257,
15389
+ "error_rate_at_0_9": 0.10553633217993079,
15390
+ "confidence_bias": 0.11803573910263576,
15391
  "top_bins": {
15392
  "0.9": {
15393
+ "n": 578,
15394
+ "errors": 61,
15395
+ "error_rate": 0.10553633217993079
15396
  },
15397
  "0.95": {
15398
+ "n": 491,
15399
+ "errors": 39,
15400
+ "error_rate": 0.07942973523421588
15401
  },
15402
  "0.99": {
15403
+ "n": 314,
15404
+ "errors": 11,
15405
+ "error_rate": 0.03503184713375796
15406
  }
15407
  },
15408
  "selective": {
15409
  "0.5": {
15410
  "coverage": 0.5,
15411
+ "accuracy": 0.9101338432122371,
15412
+ "confidence_cutoff": 0.9301925551074537
15413
  },
15414
  "0.8": {
15415
  "coverage": 0.8001912045889101,
15416
+ "accuracy": 0.7921146953405018,
15417
+ "confidence_cutoff": 0.5957019462220806
15418
  }
15419
  },
15420
  "score_mae": 0.35059676214309027,
 
66859
  }
66860
  },
66861
  "paired_comparisons": {
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
66862
  "Sol minus kev-0.8b": {
66863
  "old_core": {
66864
  "delta": 0.1360863095238095,
 
67497
  ]
67498
  }
67499
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
67500
  "Lux minus Sol": {
67501
  "old_core": {
67502
  "delta": 0.09464285714285714,
 
68134
  0.11044819701768911
68135
  ]
68136
  }
68137
+ },
68138
+ "Nox minus kev-4b": {
68139
+ "old_core": {
68140
+ "delta": 0.11097470238095238,
68141
+ "exact_fraction": "2983/26880",
68142
+ "paired_ci95": [
68143
+ 0.068154761904762,
68144
+ 0.1539462425595238
68145
+ ]
68146
+ },
68147
+ "v3_core": {
68148
+ "delta": 0.0325,
68149
+ "exact_fraction": "13/400",
68150
+ "paired_ci95": [
68151
+ 0.0033333333333334103,
68152
+ 0.06166666666666676
68153
+ ]
68154
+ },
68155
+ "v4": {
68156
+ "delta": -0.028125,
68157
+ "exact_fraction": "-9/320",
68158
+ "paired_ci95": [
68159
+ -0.062443225931676956,
68160
+ 0.00613059214504551
68161
+ ]
68162
+ },
68163
+ "v5": {
68164
+ "delta": 0.016666666666666666,
68165
+ "exact_fraction": "1/60",
68166
+ "paired_ci95": [
68167
+ -0.014652826016486136,
68168
+ 0.04855467675421749
68169
+ ]
68170
+ },
68171
+ "transfer_v9_test": {
68172
+ "delta": -0.06500956022944551,
68173
+ "exact_fraction": "-34/523",
68174
+ "paired_ci95": [
68175
+ -0.09099616858237547,
68176
+ -0.03925188928140224
68177
+ ]
68178
+ },
68179
+ "original_four_panel_mean": {
68180
+ "delta": 0.03300409226190476,
68181
+ "exact_fraction": "17743/537600",
68182
+ "paired_ci95": [
68183
+ 0.0159171480237136,
68184
+ 0.05071287650329332
68185
+ ]
68186
+ },
68187
+ "weighted_mean": {
68188
+ "delta": 0.029947226679868887,
68189
+ "exact_fraction": "1403351/46860800",
68190
+ "paired_ci95": [
68191
+ 0.013201872388473554,
68192
+ 0.047038737697081195
68193
+ ]
68194
+ }
68195
+ },
68196
+ "Nox minus Qwen3.5-4B": {
68197
+ "old_core": {
68198
+ "delta": 0.13113839285714285,
68199
+ "exact_fraction": "235/1792",
68200
+ "paired_ci95": [
68201
+ 0.09270647321428567,
68202
+ 0.16990420386904764
68203
+ ]
68204
+ },
68205
+ "v3_core": {
68206
+ "delta": 0.08458333333333333,
68207
+ "exact_fraction": "203/2400",
68208
+ "paired_ci95": [
68209
+ 0.05416666666666664,
68210
+ 0.11541666666666661
68211
+ ]
68212
+ },
68213
+ "v4": {
68214
+ "delta": -0.0890625,
68215
+ "exact_fraction": "-57/640",
68216
+ "paired_ci95": [
68217
+ -0.1342959349117049,
68218
+ -0.045753734276729595
68219
+ ]
68220
+ },
68221
+ "v5": {
68222
+ "delta": 0.06458333333333334,
68223
+ "exact_fraction": "31/480",
68224
+ "paired_ci95": [
68225
+ 0.028211805555555566,
68226
+ 0.10191039756413821
68227
+ ]
68228
+ },
68229
+ "transfer_v9_test": {
68230
+ "delta": 0.0076481835564053535,
68231
+ "exact_fraction": "4/523",
68232
+ "paired_ci95": [
68233
+ -0.022396190701526143,
68234
+ 0.03838776086034498
68235
+ ]
68236
+ },
68237
+ "original_four_panel_mean": {
68238
+ "delta": 0.04781063988095238,
68239
+ "exact_fraction": "25703/537600",
68240
+ "paired_ci95": [
68241
+ 0.028947119327501308,
68242
+ 0.06683696978735856
68243
+ ]
68244
+ },
68245
+ "weighted_mean": {
68246
+ "delta": 0.05796270372393699,
68247
+ "exact_fraction": "1018567/17572800",
68248
+ "paired_ci95": [
68249
+ 0.04083969253575949,
68250
+ 0.07488554488575702
68251
+ ]
68252
+ }
68253
+ },
68254
+ "Nox minus Decider": {
68255
+ "old_core": {
68256
+ "delta": 0.18988095238095237,
68257
+ "exact_fraction": "319/1680",
68258
+ "paired_ci95": [
68259
+ 0.1443443080357141,
68260
+ 0.23571428571428554
68261
+ ]
68262
+ },
68263
+ "v3_core": {
68264
+ "delta": 0.052083333333333336,
68265
+ "exact_fraction": "5/96",
68266
+ "paired_ci95": [
68267
+ 0.013333333333333308,
68268
+ 0.08999999999999997
68269
+ ]
68270
+ },
68271
+ "v4": {
68272
+ "delta": -0.1296875,
68273
+ "exact_fraction": "-83/640",
68274
+ "paired_ci95": [
68275
+ -0.17179976851851855,
68276
+ -0.08906249999999993
68277
+ ]
68278
+ },
68279
+ "v5": {
68280
+ "delta": 0.01875,
68281
+ "exact_fraction": "3/160",
68282
+ "paired_ci95": [
68283
+ -0.012535333369549428,
68284
+ 0.04969926075268807
68285
+ ]
68286
+ },
68287
+ "transfer_v9_test": {
68288
+ "delta": 0.0028680688336520078,
68289
+ "exact_fraction": "3/1046",
68290
+ "paired_ci95": [
68291
+ -0.030160845668367766,
68292
+ 0.03653934071222337
68293
+ ]
68294
+ },
68295
+ "original_four_panel_mean": {
68296
+ "delta": 0.03275669642857143,
68297
+ "exact_fraction": "587/17920",
68298
+ "paired_ci95": [
68299
+ 0.013155496386521998,
68300
+ 0.05228677715137045
68301
+ ]
68302
+ },
68303
+ "weighted_mean": {
68304
+ "delta": 0.05377470437266685,
68305
+ "exact_fraction": "7559777/140582400",
68306
+ "paired_ci95": [
68307
+ 0.03467000712516513,
68308
+ 0.07268914300878532
68309
+ ]
68310
+ }
68311
  }
68312
  },
68313
  "method": {
 
68335
  "transfer_accuracy": "Upstream ordered probability argmax, clean knowable records, micro average. Refused/missing probabilities count wrong; unknown evidence is not accuracy.",
68336
  "transfer_clusters": "Union IDs/group/parent/control/pair plus exact state SHA and ordered-request IDs, built over all variants then scored only clean knowable rows. Complete components jointly resampled; variable sampled denominator. No whole-family grouping.",
68337
  "pairing": "Identical sampled source components across all models; five panels sampled independently. Original four marginal CIs preserved.",
68338
+ "accuracy_arithmetic": "Exact rational integer correct/requested; weights not fitted to outcomes.",
68339
  "interval_role": "Descriptive paired95% intervals; no positive-CI release requirement invented.",
68340
  "diagnostics": "Full per-task/proper-score/order/unknown-evidence results remain separate; no accuracy/latency/ECE unit mixing.",
68341
+ "exposure": "Observed regression tests; weighting user-approved after earlier results; not a blind prospective benchmark."
68342
  },
68343
  "prior_weighted_scores": {
68344
  "Jev": {
 
68358
  ]
68359
  },
68360
  "Nox": {
68361
+ "point": 0.724150437750387,
68362
+ "exact_fraction": "203605613/281164800",
68363
+ "interval_status": "Not recomputed for alternate weighting; only the point is displayed."
 
 
 
68364
  },
68365
  "kev-9b": {
68366
  "point": 0.7201455089684057,
 
68445
  },
68446
  "hosted_frontier": "All official calls returned jev-1.13.0; closed weight identity/input retention cannot be independently inspected.",
68447
  "size_note": "Qwen3.5 family size labels name their official parent checkpoint. Lux deployed text+decision weights total7,940,895,744 parameters; no vision or generativeLMhead in its serving bundle.",
68448
+ "presentation_note": "Latest released result per model. Nox uses its qualified Choice null-description rendering; other displayed model results are unchanged. No historical model rows.",
68449
+ "update_kind": "Nox SystemOne Choice semantics: null descriptions use their key text; model weights, tokenizer and temperature unchanged."
68450
  }
metrics/evaluation-provenance.json CHANGED
@@ -1,145 +1,148 @@
1
  {
2
- "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
3
- "manifest_sha256": "3199a40b2719cad9b1e815629cf5ebad923a9a399188c1199a9d27e918bda9e3",
 
4
  "qualified_runtime": {
5
  "python": "3.12.13",
6
  "numpy": "2.3.5"
7
  },
8
- "source_sha256": "b00c60f21ace9bd258a1a79673b8cdbce3101ada2fed5547811f70bf030fcd91",
9
  "model_manifest": {
10
- "Nox": {
11
  "panels": {
12
  "old_core": {
13
- "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/old_core/normalized.jsonl",
14
- "sha256": "9b6b7db907e8e458cd33982c0f7c45e5052c3248285d927990f160d5570c8976"
 
15
  },
16
  "v3_core": {
17
- "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v3_core/normalized.jsonl",
18
- "sha256": "49575aef2261c191e289d781857664c57fd6181701c03f02515292f2fa31b40c"
 
19
  },
20
  "v4": {
21
- "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v4/normalized.jsonl",
22
- "sha256": "9d85618b1a05c2ed68f6a188dadf19a64d6ede6d5a1b0ef58afa27a13906238b"
 
23
  },
24
  "v5": {
25
- "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v5/normalized.jsonl",
26
- "sha256": "5eb7faaa45a22dad414af1363d6087b8c379c86599df3c14483a3a648ca6bfd5"
 
27
  }
28
  },
29
  "transfer_rows": {
30
- "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Nox/v9-transfer-v9-test-published-temperature-rows.json",
31
- "sha256": "77b92a18433b260ac0e6f4e0b6c9ec3a94e16aaed870a4078b1d7f32ef18b169"
32
  },
33
  "evidence": [
34
  {
35
- "path": "results/expanded/nox-retention-v1/QUALITY-COLLECTION.json",
36
- "sha256": "7608258282ec1e9731e4d2ea4dee7206fda8900055dc64f9193d2566172ca86a"
 
 
 
 
37
  },
38
  {
39
- "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Nox/REPORT.json",
40
- "sha256": "8e78910746c9e78862b3aadd94490054a681f3ad013b43ec19bfbc2434111ed7"
41
  }
42
  ]
43
  },
44
- "Sol": {
45
  "panels": {
46
  "old_core": {
47
- "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/old_core/normalized.jsonl",
48
- "sha256": "5fc0b3e49a5de576f7621839130855c178b12d942dc357aaf6b4dd9735e0a6a8"
49
  },
50
  "v3_core": {
51
- "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v3_core/normalized.jsonl",
52
- "sha256": "b2c99fbd739b849a66a55929a658de4517a7ffaa4496255566e6972ba28d4f15"
53
  },
54
  "v4": {
55
- "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v4/normalized.jsonl",
56
- "sha256": "3c28302750e02d4cead0aacd66a323734e773f726c4c41db658dd6461251b28b"
57
  },
58
  "v5": {
59
- "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v5/normalized.jsonl",
60
- "sha256": "6ddb0496012b596ea122afa4db4c4158e57dfb934e9ae29eb006756b652e31e1"
61
  }
62
  },
63
  "transfer_rows": {
64
- "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/v9-transfer-v9-test-published-temperature-rows.json",
65
- "sha256": "b05799417cb3715de633135c58f713960b75bdb9ed30eb8912a197d33ee3d85b"
66
  },
67
  "evidence": [
68
  {
69
- "path": "results/expanded/sol-composition-v2/QUALITY-COLLECTION.json",
70
- "sha256": "9b77aac87b85cbf3a74e86af4a2b9a03fd9e1ed854356a92ff54ccd0d47a16f4"
71
  },
72
  {
73
- "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/REPORT.json",
74
- "sha256": "5d8ebb339293687cec9089f54fb96aabf3330f16111487a49489f5e52bbb6f29"
 
 
 
 
 
 
 
 
 
 
 
 
75
  }
76
  ]
77
  },
78
- "kev-0.8b": {
79
  "panels": {
80
  "old_core": {
81
- "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-old_core-normalized.jsonl",
82
- "sha256": "b492b4512ac164f83a319f31bbc9e086cae12f87ffe2c14f307481ffb240807c"
83
  },
84
  "v3_core": {
85
- "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v3_core-normalized.jsonl",
86
- "sha256": "70cbfc69dfc72ac0c5ae01305ec63a9f31129c115b7068c6824cd22ead013d68"
87
  },
88
  "v4": {
89
- "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v4-normalized.jsonl",
90
- "sha256": "761dfd69acbface101f2c21fd27d2db684c5c8b8cc425890a075729ae1f7db24"
91
  },
92
  "v5": {
93
- "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v5-normalized.jsonl",
94
- "sha256": "51b6fc5f826ca8db44df43885f34d857ddd5b5a66453bbdc1a83fcf488928bdb"
95
  }
96
  },
97
  "transfer_rows": {
98
- "path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/v9-transfer-v9-test-published-temperature-rows.json",
99
- "sha256": "d6abed88001aeac630219466f02c1aad89fcbd354b3ae533507f3c2f7fe42898"
100
  },
101
  "evidence": [
102
  {
103
- "path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/REPORT.json",
104
- "sha256": "6760dcbaa7d85e6c2e893e4c80ab8e66e0b84d7cfa26ae9e25909f48a05f8de4"
105
  },
106
  {
107
- "path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
108
- "sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
109
- }
110
- ]
111
- },
112
- "kev-4b": {
113
- "panels": {
114
- "old_core": {
115
- "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-old_core-normalized.jsonl",
116
- "sha256": "41dfbb7c9c4381effc219f9123b51bd4d688b3e746cec974ef930730b9dfd6b1"
117
  },
118
- "v3_core": {
119
- "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v3_core-normalized.jsonl",
120
- "sha256": "2ec54da92c4e72412e89b3c07013b54729f5899ba8542eeadf962b04a1f3a4b8"
121
  },
122
- "v4": {
123
- "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v4-normalized.jsonl",
124
- "sha256": "88c576c3d62ec11686dd5ebf360971a9087779e45a346897e2badf874a754a12"
125
  },
126
- "v5": {
127
- "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v5-normalized.jsonl",
128
- "sha256": "d33bbb60134314dcb550060176ed0a3ae91ec0c6b033cceb40e3255b33666898"
129
- }
130
- },
131
- "transfer_rows": {
132
- "path": "analysis/kev-reciprocal-v1/reports/kev-4b/v9-transfer-v9-test-published-temperature-rows.json",
133
- "sha256": "c2cb16a936be4c10ca6b1870c69426ddd6b51629c401f71ae1da652b3bd2ac7e"
134
- },
135
- "evidence": [
136
  {
137
- "path": "analysis/kev-reciprocal-v1/reports/kev-4b/REPORT.json",
138
- "sha256": "366fb0d27a4d732e5e270c264521d9a93ccabc4f0ddf5b7d3e9f1ecc5f460406"
139
  },
140
  {
141
- "path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
142
- "sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
143
  }
144
  ]
145
  },
@@ -177,129 +180,37 @@
177
  }
178
  ]
179
  },
180
- "Jev": {
181
- "panels": {
182
- "old_core": {
183
- "path": "eval/heldout/official-predictions.jsonl",
184
- "sha256": "069e48e1fdf1b5f3d95cc0097e4179ac3faab97bb357b0ce897cfb463b87cbaf",
185
- "bytes": 185263
186
- },
187
- "v3_core": {
188
- "path": "eval/heldout/v3/official/normalized/core.jsonl",
189
- "sha256": "54707354bfe5d96b5de56e6dc1b38cba0a610eea516b28fbff013b94f422b397",
190
- "bytes": 368716
191
- },
192
- "v4": {
193
- "path": "eval/heldout/v4/official/normalized/core.jsonl",
194
- "sha256": "d6583f903e974d70747fae2944014828954dc4d7344ceddcbb2cd1b06fbfc799",
195
- "bytes": 183671
196
- },
197
- "v5": {
198
- "path": "eval/heldout/v5/official/normalized/core.jsonl",
199
- "sha256": "f883164c4c83db8e42ebe03f7e165afbd2f81f2b541bcf8008a2cc06a60e1ee7",
200
- "bytes": 200316
201
- }
202
- },
203
- "transfer_rows": {
204
- "path": "analysis/decision-benchmark-v3-official/scored/ROWS.json",
205
- "sha256": "be29a7f3804ad84993ffb06f05ada7e2d8ed25317504c4fc17fe5dc306fd788c"
206
- },
207
- "evidence": [
208
- {
209
- "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
210
- "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
211
- },
212
- {
213
- "path": "analysis/decision-benchmark-v3-official/scored/REPORT.json",
214
- "sha256": "2fdbaf7db1e4b877762e33a87b23a236a065eaa3078b14c2d7d579ad44f9ea24"
215
- },
216
- {
217
- "path": "analysis/decision-benchmark-v3-official/INDEPENDENT-PROJECTION-REVIEW.json",
218
- "sha256": "b9506539e2c4a863aef16f7dab0517c175e85863416bc622c60df05489a9eb52"
219
- }
220
- ]
221
- },
222
- "Laya-base": {
223
  "panels": {
224
- "v4": {
225
- "path": "results/public-roster-baselines-v1/laya-english/worker/v4/normalized.jsonl",
226
- "sha256": "31da1d4140da69d2677172707eb301843cf8fdfb43200cba93be910321f922ba"
227
- },
228
- "v5": {
229
- "path": "results/public-roster-baselines-v1/laya-english/worker/v5/normalized.jsonl",
230
- "sha256": "159614ceede814b578e1e025836a595c8c71a9bec4896caf6454d7bf54c4c44d"
231
- },
232
  "old_core": {
233
- "path": "eval/heldout/primary-v2/laya-english/core/normalized.jsonl",
234
- "sha256": "850643d2daeb47a6909985003736cab0f1f0e5a87631a35e27192f0733efd756"
235
  },
236
  "v3_core": {
237
- "path": "eval/heldout/v3/baselines/laya-english/core/normalized.jsonl",
238
- "sha256": "dec19caeb8074e25289e43a367392172da21f0c2ccdb403a059e2e4a2364a64b"
239
- }
240
- },
241
- "transfer_rows": {
242
- "path": "analysis/decision-benchmark-v3-baselines/laya-english/ROWS.json",
243
- "sha256": "19138909259400facb5d3d5402d0d736f870695fe3ee46019f2e438088c57275"
244
- },
245
- "evidence": [
246
- {
247
- "path": "results/public-roster-baselines-v1/laya-english/worker/COMPLETE.json",
248
- "sha256": "cef0da797ae14cebaf1e4b9edf9b16c3cd9a412a30435d7d2517b8ba6dbe8b98"
249
- },
250
- {
251
- "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
252
- "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
253
- },
254
- {
255
- "path": "analysis/decision-benchmark-v3-baselines/laya-english/REPORT.json",
256
- "sha256": "cb5935bb399dcf23ffe24f855452b2e074bcbb3d023cfc7dd439342e54f02c6b"
257
  },
258
- {
259
- "path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
260
- "sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
261
- }
262
- ]
263
- },
264
- "Laya-multilingual": {
265
- "panels": {
266
  "v4": {
267
- "path": "results/public-roster-baselines-v1/laya-multilingual/worker/v4/normalized.jsonl",
268
- "sha256": "814c1aa1f4d400942984ecb3786d3fd5fd15286d9b767a519c91993577795290"
269
  },
270
  "v5": {
271
- "path": "results/public-roster-baselines-v1/laya-multilingual/worker/v5/normalized.jsonl",
272
- "sha256": "6d4dbc16c7b977b15895e0e313dabc6d9db43df75ec56e89373d91f4f13ec676"
273
- },
274
- "old_core": {
275
- "path": "eval/heldout/primary-v2/laya-multilingual/core/normalized.jsonl",
276
- "sha256": "9ea39841c4d5901cd9ab860c980d2f3b55358be289c7e01b9b47982927bf849d"
277
- },
278
- "v3_core": {
279
- "path": "eval/heldout/v3/baselines/laya-multilingual/core/normalized.jsonl",
280
- "sha256": "a0b7ce9daf4be2c7831bcab6b5e96c8dcb6bef133d39475f72241ab18faccdd2"
281
  }
282
  },
283
  "transfer_rows": {
284
- "path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/ROWS.json",
285
- "sha256": "d32db6cf6576f287210d31336456497f9d6926179366db29304659baf9783b43"
286
  },
287
  "evidence": [
288
  {
289
- "path": "results/public-roster-baselines-v1/laya-multilingual/worker/COMPLETE.json",
290
- "sha256": "b190d683fc4c4762c2805a03d52628a53e125d228a3105c8ac2ddda00a5364e4"
291
- },
292
- {
293
- "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
294
- "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
295
- },
296
- {
297
- "path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/REPORT.json",
298
- "sha256": "d20d3d58f76ca2758545f83d358abd747f86d581c897d1fb521371d1000d6154"
299
  },
300
  {
301
- "path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
302
- "sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
303
  }
304
  ]
305
  },
@@ -361,99 +272,171 @@
361
  }
362
  ]
363
  },
364
- "Lux": {
365
  "panels": {
366
  "old_core": {
367
- "path": "results/lux-new-host-v1/quality/lux/old_core/normalized.jsonl",
368
- "sha256": "a659a7c7fef2871bbf806e842ea8ca81f645dcf610f0f2f6622d8d32e76a0836"
369
  },
370
  "v3_core": {
371
- "path": "results/lux-new-host-v1/quality/lux/v3_core/normalized.jsonl",
372
- "sha256": "da8903a07e36a9687f6ee1d0fd8dbc684e61c49ef5fea30c346e46bff5a90e94"
373
  },
374
  "v4": {
375
- "path": "results/lux-new-host-v1/quality/lux/v4/normalized.jsonl",
376
- "sha256": "26cd7482693ec1ffa1a93a156abfc88378e4872e320f3b9b775763635086d488"
377
  },
378
  "v5": {
379
- "path": "results/lux-new-host-v1/quality/lux/v5/normalized.jsonl",
380
- "sha256": "759c773e629a16317271de850780a8734533c795046122f850914c4482d1e72b"
381
  }
382
  },
383
  "transfer_rows": {
384
- "path": "analysis/decision-benchmark-v3-baselines/Lux/ROWS.json",
385
- "sha256": "bafb55fc86762061e537382360eda14a8a8f8c5e4e111c2fda1f82cd6c06793f"
386
  },
387
  "evidence": [
388
  {
389
- "path": "results/lux-new-host-v1/quality/lux/COMPLETE.json",
390
- "sha256": "9600ff792cbda13af5505c804cf63af94d2ee895f2912f0645a7d493ae1b7a9c"
391
  },
392
  {
393
- "path": "results/lux-new-host-v1/quality/lux/metadata.json",
394
- "sha256": "0d4509fae7eb3630118eaebc24e75a7b02f717d9a0ad74811fa27358c33c9cc5"
395
  },
396
  {
397
- "path": "analysis/decision-benchmark-v3-baselines/Lux-projected/COMPLETE.json",
398
- "sha256": "f58d02df78f297d4feebe39facfac072951111d8d6a19e2c1069f1b90250c622"
399
  },
400
  {
401
- "path": "analysis/decision-benchmark-v3-baselines/Lux/REPORT.json",
402
- "sha256": "f31703c342efe043a1e1534da189808bb7e0de0641fbaf3f91faa4a1f928a299"
403
  },
404
  {
405
- "path": "analysis/decoder4b/lux9b-training-v1/CHECKPOINT-HELDOUT-COLLECTED.json",
406
- "sha256": "d1c42515ab225b27ea011a6cf2262fe12dd7001e3b4d2cb64a23e7642c7a0d91"
 
 
 
 
407
  }
408
  ]
409
  },
410
- "Decider": {
411
  "panels": {
412
  "old_core": {
413
- "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/decider/quality/normalized.jsonl",
414
- "sha256": "e1b14f0f9ec521bffbde5945684bc2410a413b6b6418d65ac4fc702a7b5e7937"
415
  },
416
  "v3_core": {
417
- "path": "eval/heldout/v3/baselines/decider/core/normalized.jsonl",
418
- "sha256": "12a1f554decf1aff608743e7b4a44681289ea084ae25f383e519b1f34917d46f"
419
  },
420
  "v4": {
421
- "path": "eval/heldout/v4/baselines/decider/core/normalized.jsonl",
422
- "sha256": "626ef366ae471ec10dfb89ef2ff2f7b29a6c879976d3ff14fcbb2bafd7d9b045"
423
  },
424
  "v5": {
425
- "path": "eval/heldout/v5/open-baselines-v1/decider/core/normalized.jsonl",
426
- "sha256": "ae16b17ac8e4208940c0c3e9e04f25e94686b5d16bf46f64f53669baf41aeb71"
427
  }
428
  },
429
  "transfer_rows": {
430
- "path": "analysis/decision-benchmark-v3-baselines/decider/ROWS.json",
431
- "sha256": "6da5b5692a8a0cea3dc8f6d02d911fafd9bbf62e286a9be8db27c89f77a897e4"
432
  },
433
  "evidence": [
434
  {
435
- "path": "results/accelerated-transfer-baselines-v1/decider/worker/COMPLETE.json",
436
- "sha256": "e1ee0df8df59ec3eb2b49ba947a8508994a24bce226b7f2cccbe25f6835b99af"
437
  },
438
  {
439
- "path": "results/accelerated-transfer-baselines-v1/decider/worker/METADATA.json",
440
- "sha256": "f00061c1b0bb642bf1353aae867d19df988c01b0428b58206b31249521bcbff8"
441
  },
442
  {
443
- "path": "results/accelerated-transfer-baselines-v1/decider/worker/predictions.jsonl",
444
- "sha256": "f69e87fa65b09289e6861c1d789270aea5a92d23ac47817342647ed08f406df7"
445
  },
446
  {
447
- "path": "analysis/decision-benchmark-v3-baselines/decider/REPORT.json",
448
- "sha256": "5aae029f34524dfebc8e79128da8c39229c5747461dfce07ca59956da4416951"
449
  },
450
  {
451
  "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
452
  "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
453
  },
454
  {
455
- "path": "analysis/decision-benchmark-v3-baselines/decider/COMPOSABLE-MANIFEST.json",
456
- "sha256": "a49848e86feafaa4f36a33287910035e7605e344791aae139305b78cb304a77f"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
457
  }
458
  ]
459
  },
@@ -507,60 +490,93 @@
507
  }
508
  ]
509
  },
510
- "Qwen3.5-4B": {
511
  "panels": {
512
- "old_core": {
513
- "path": "eval/heldout/primary-v2/base-4b/core/normalized.jsonl",
514
- "sha256": "d20ad845066ff81caaf60ebee44df941872bdc642da78b5f3b3da4a33dd9a08d"
515
- },
516
- "v3_core": {
517
- "path": "eval/heldout/v3/baselines/base-4b/core/normalized.jsonl",
518
- "sha256": "7541893d2b2ec0ec2ffc513f50c568eaef860b65bcef9a1747996c116fff9661"
519
- },
520
  "v4": {
521
- "path": "eval/heldout/v4/baselines/base-4b/core/normalized.jsonl",
522
- "sha256": "7a6d252ccaa72bf54b94445190bbad023fb465a3eaf4cf657532e2cdd4f750f5"
523
  },
524
  "v5": {
525
- "path": "eval/heldout/v5/open-baselines-v1/base-4b/core/normalized.jsonl",
526
- "sha256": "c01ad5daed65267e9aec40ed9e97ece1b6511e4cab17ebff09339b41524fb7f1"
 
 
 
 
 
 
 
 
527
  }
528
  },
529
  "transfer_rows": {
530
- "path": "analysis/decision-benchmark-v3-baselines/base-4b/ROWS.json",
531
- "sha256": "37ab17d94d640ee46283e1f23723b9ea05133c22a8a9b62a97c756a89c5de785"
532
  },
533
  "evidence": [
534
  {
535
- "path": "results/remaining-transfer-baselines-v1/base-4b/worker/COMPLETE.json",
536
- "sha256": "72abc4f87a9234f95eb3b0a72a9e7d66798779d5982c757faffb610003b1b242"
537
  },
538
  {
539
- "path": "results/remaining-transfer-baselines-v1/base-4b/worker/METADATA.json",
540
- "sha256": "871fed92cf762187015e59174a042808dbf3fc88f7791aa0ddd530ae923196ab"
541
  },
542
  {
543
- "path": "results/remaining-transfer-baselines-v1/base-4b/worker/predictions.jsonl",
544
- "sha256": "8c47edda469f8914eccc2ecafa91c245752812e830cb9f29445653102522854a"
545
  },
546
  {
547
- "path": "analysis/decision-benchmark-v3-baselines/base-4b/REPORT.json",
548
- "sha256": "c2a0aafbe2b7c85b3c07a4b9843182a3361654f57883311155aa3b05cfc7002d"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
549
  },
 
 
 
 
 
 
 
 
 
 
550
  {
551
- "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
552
- "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
553
  },
554
  {
555
- "path": "analysis/decision-benchmark-v3-baselines/base-4b/COMPOSABLE-MANIFEST.json",
556
- "sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
 
 
 
 
 
 
 
 
557
  }
558
  ]
559
  }
560
  },
561
  "task_rows": 54,
562
  "rank_public_models": [
563
- "Jev",
564
  "Lux",
565
  "Nox",
566
  "kev-9b",
@@ -572,7 +588,8 @@
572
  "kev-0.8b",
573
  "Qwen3.5-2B",
574
  "Laya-base",
575
- "Laya-multilingual"
 
576
  ],
577
  "internal_extra_models_not_public_rank": [
578
  "llm2jev-2b",
@@ -655,7 +672,5 @@
655
  }
656
  },
657
  "source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
658
- },
659
- "frozen_protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
660
- "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
661
  }
 
1
  {
2
+ "frozen_protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
3
+ "statistics_sha256": "b319a993ac569ea1fa8da027ccb2195868557e4efe232a5d06091db2884086be",
4
+ "manifest_sha256": "cf455486f8eba2b7158bbb4b5bcdcd8a4bd4daee4f05184b1df61b5f70ef2b04",
5
  "qualified_runtime": {
6
  "python": "3.12.13",
7
  "numpy": "2.3.5"
8
  },
9
+ "source_sha256": "f2f72fe161944599b1114df29ff095786bf661cca59afd9d784b7ca6d1430c99",
10
  "model_manifest": {
11
+ "Jev": {
12
  "panels": {
13
  "old_core": {
14
+ "path": "eval/heldout/official-predictions.jsonl",
15
+ "sha256": "069e48e1fdf1b5f3d95cc0097e4179ac3faab97bb357b0ce897cfb463b87cbaf",
16
+ "bytes": 185263
17
  },
18
  "v3_core": {
19
+ "path": "eval/heldout/v3/official/normalized/core.jsonl",
20
+ "sha256": "54707354bfe5d96b5de56e6dc1b38cba0a610eea516b28fbff013b94f422b397",
21
+ "bytes": 368716
22
  },
23
  "v4": {
24
+ "path": "eval/heldout/v4/official/normalized/core.jsonl",
25
+ "sha256": "d6583f903e974d70747fae2944014828954dc4d7344ceddcbb2cd1b06fbfc799",
26
+ "bytes": 183671
27
  },
28
  "v5": {
29
+ "path": "eval/heldout/v5/official/normalized/core.jsonl",
30
+ "sha256": "f883164c4c83db8e42ebe03f7e165afbd2f81f2b541bcf8008a2cc06a60e1ee7",
31
+ "bytes": 200316
32
  }
33
  },
34
  "transfer_rows": {
35
+ "path": "analysis/decision-benchmark-v3-official/scored/ROWS.json",
36
+ "sha256": "be29a7f3804ad84993ffb06f05ada7e2d8ed25317504c4fc17fe5dc306fd788c"
37
  },
38
  "evidence": [
39
  {
40
+ "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
41
+ "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
42
+ },
43
+ {
44
+ "path": "analysis/decision-benchmark-v3-official/scored/REPORT.json",
45
+ "sha256": "2fdbaf7db1e4b877762e33a87b23a236a065eaa3078b14c2d7d579ad44f9ea24"
46
  },
47
  {
48
+ "path": "analysis/decision-benchmark-v3-official/INDEPENDENT-PROJECTION-REVIEW.json",
49
+ "sha256": "b9506539e2c4a863aef16f7dab0517c175e85863416bc622c60df05489a9eb52"
50
  }
51
  ]
52
  },
53
+ "Lux": {
54
  "panels": {
55
  "old_core": {
56
+ "path": "results/lux-new-host-v1/quality/lux/old_core/normalized.jsonl",
57
+ "sha256": "a659a7c7fef2871bbf806e842ea8ca81f645dcf610f0f2f6622d8d32e76a0836"
58
  },
59
  "v3_core": {
60
+ "path": "results/lux-new-host-v1/quality/lux/v3_core/normalized.jsonl",
61
+ "sha256": "da8903a07e36a9687f6ee1d0fd8dbc684e61c49ef5fea30c346e46bff5a90e94"
62
  },
63
  "v4": {
64
+ "path": "results/lux-new-host-v1/quality/lux/v4/normalized.jsonl",
65
+ "sha256": "26cd7482693ec1ffa1a93a156abfc88378e4872e320f3b9b775763635086d488"
66
  },
67
  "v5": {
68
+ "path": "results/lux-new-host-v1/quality/lux/v5/normalized.jsonl",
69
+ "sha256": "759c773e629a16317271de850780a8734533c795046122f850914c4482d1e72b"
70
  }
71
  },
72
  "transfer_rows": {
73
+ "path": "analysis/decision-benchmark-v3-baselines/Lux/ROWS.json",
74
+ "sha256": "bafb55fc86762061e537382360eda14a8a8f8c5e4e111c2fda1f82cd6c06793f"
75
  },
76
  "evidence": [
77
  {
78
+ "path": "results/lux-new-host-v1/quality/lux/COMPLETE.json",
79
+ "sha256": "9600ff792cbda13af5505c804cf63af94d2ee895f2912f0645a7d493ae1b7a9c"
80
  },
81
  {
82
+ "path": "results/lux-new-host-v1/quality/lux/metadata.json",
83
+ "sha256": "0d4509fae7eb3630118eaebc24e75a7b02f717d9a0ad74811fa27358c33c9cc5"
84
+ },
85
+ {
86
+ "path": "analysis/decision-benchmark-v3-baselines/Lux-projected/COMPLETE.json",
87
+ "sha256": "f58d02df78f297d4feebe39facfac072951111d8d6a19e2c1069f1b90250c622"
88
+ },
89
+ {
90
+ "path": "analysis/decision-benchmark-v3-baselines/Lux/REPORT.json",
91
+ "sha256": "f31703c342efe043a1e1534da189808bb7e0de0641fbaf3f91faa4a1f928a299"
92
+ },
93
+ {
94
+ "path": "analysis/decoder4b/lux9b-training-v1/CHECKPOINT-HELDOUT-COLLECTED.json",
95
+ "sha256": "d1c42515ab225b27ea011a6cf2262fe12dd7001e3b4d2cb64a23e7642c7a0d91"
96
  }
97
  ]
98
  },
99
+ "Nox": {
100
  "panels": {
101
  "old_core": {
102
+ "path": "results/nox-null-description-v1/full-regression/old_core/normalized.jsonl",
103
+ "sha256": "dbd18273dbbb9d09ba37d19dfedb0fcd6b64841d528632f04d39fd7955799967"
104
  },
105
  "v3_core": {
106
+ "path": "results/nox-null-description-v1/full-regression/v3_core/normalized.jsonl",
107
+ "sha256": "5286d1c3a78649e3284a018b1518f94cf32d8f0ad43a49c6a14c993524a3aff8"
108
  },
109
  "v4": {
110
+ "path": "results/nox-null-description-v1/full-regression/v4/normalized.jsonl",
111
+ "sha256": "077b752f4866bd0fc83db9b31c6227fd5015ce3585d2138c1a7a803c433d472b"
112
  },
113
  "v5": {
114
+ "path": "results/nox-null-description-v1/full-regression/v5/normalized.jsonl",
115
+ "sha256": "d57f07df59ff9b0e6ace0658e7917746e2b8d3bf518309a3858045cd175d79df"
116
  }
117
  },
118
  "transfer_rows": {
119
+ "path": "analysis/nox-null-description-candidate-v4/transfer/ROWS.json",
120
+ "sha256": "7cee9e036ee53a3c41918cfafc73ad0d68f16b266abcae43077d1d69247afdfe"
121
  },
122
  "evidence": [
123
  {
124
+ "path": "results/nox-null-description-v1/COMPLETE.json",
125
+ "sha256": "b813aa5db8c43b728b94abbaef20cefc8d3d84b51f7873ca46b3afb59aab6fe4"
126
  },
127
  {
128
+ "path": "results/nox-null-description-v1/full-regression/COMPLETE.json",
129
+ "sha256": "ad80e5bfb2558ee3b8fb1fc80617fcefbc1fbb8e58b8155e3cf5dfde38528754"
 
 
 
 
 
 
 
 
130
  },
131
+ {
132
+ "path": "results/nox-null-description-v1/transfer-v9/COMPLETE.json",
133
+ "sha256": "93a8767f7245b93764e527e616f236b5fa46b24a087a7e8ff66d35eee9f9fd0f"
134
  },
135
+ {
136
+ "path": "results/nox-null-description-v1/full-regression-ACTUAL-EXIT.json",
137
+ "sha256": "dca6e05615c9c82a2d506438fb4e7c9abfe03053661515ffc9a4c5b176fd3ed8"
138
  },
 
 
 
 
 
 
 
 
 
 
139
  {
140
+ "path": "results/nox-null-description-v1/transfer-v9-ACTUAL-EXIT.json",
141
+ "sha256": "09c4c05549253c386f72b52e8ad4fbbc47d05af59715caebf9d01a6e8734d0af"
142
  },
143
  {
144
+ "path": "analysis/nox-null-description-candidate-v4/transfer/REPORT.json",
145
+ "sha256": "f02e447787c0c5afcb8acbd06408407719ebfc191e8501e8a887cd96c150f81b"
146
  }
147
  ]
148
  },
 
180
  }
181
  ]
182
  },
183
+ "kev-4b": {
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
184
  "panels": {
 
 
 
 
 
 
 
 
185
  "old_core": {
186
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-old_core-normalized.jsonl",
187
+ "sha256": "41dfbb7c9c4381effc219f9123b51bd4d688b3e746cec974ef930730b9dfd6b1"
188
  },
189
  "v3_core": {
190
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v3_core-normalized.jsonl",
191
+ "sha256": "2ec54da92c4e72412e89b3c07013b54729f5899ba8542eeadf962b04a1f3a4b8"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
192
  },
 
 
 
 
 
 
 
 
193
  "v4": {
194
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v4-normalized.jsonl",
195
+ "sha256": "88c576c3d62ec11686dd5ebf360971a9087779e45a346897e2badf874a754a12"
196
  },
197
  "v5": {
198
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v5-normalized.jsonl",
199
+ "sha256": "d33bbb60134314dcb550060176ed0a3ae91ec0c6b033cceb40e3255b33666898"
 
 
 
 
 
 
 
 
200
  }
201
  },
202
  "transfer_rows": {
203
+ "path": "analysis/kev-reciprocal-v1/reports/kev-4b/v9-transfer-v9-test-published-temperature-rows.json",
204
+ "sha256": "c2cb16a936be4c10ca6b1870c69426ddd6b51629c401f71ae1da652b3bd2ac7e"
205
  },
206
  "evidence": [
207
  {
208
+ "path": "analysis/kev-reciprocal-v1/reports/kev-4b/REPORT.json",
209
+ "sha256": "366fb0d27a4d732e5e270c264521d9a93ccabc4f0ddf5b7d3e9f1ecc5f460406"
 
 
 
 
 
 
 
 
210
  },
211
  {
212
+ "path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
213
+ "sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
214
  }
215
  ]
216
  },
 
272
  }
273
  ]
274
  },
275
+ "Decider": {
276
  "panels": {
277
  "old_core": {
278
+ "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/decider/quality/normalized.jsonl",
279
+ "sha256": "e1b14f0f9ec521bffbde5945684bc2410a413b6b6418d65ac4fc702a7b5e7937"
280
  },
281
  "v3_core": {
282
+ "path": "eval/heldout/v3/baselines/decider/core/normalized.jsonl",
283
+ "sha256": "12a1f554decf1aff608743e7b4a44681289ea084ae25f383e519b1f34917d46f"
284
  },
285
  "v4": {
286
+ "path": "eval/heldout/v4/baselines/decider/core/normalized.jsonl",
287
+ "sha256": "626ef366ae471ec10dfb89ef2ff2f7b29a6c879976d3ff14fcbb2bafd7d9b045"
288
  },
289
  "v5": {
290
+ "path": "eval/heldout/v5/open-baselines-v1/decider/core/normalized.jsonl",
291
+ "sha256": "ae16b17ac8e4208940c0c3e9e04f25e94686b5d16bf46f64f53669baf41aeb71"
292
  }
293
  },
294
  "transfer_rows": {
295
+ "path": "analysis/decision-benchmark-v3-baselines/decider/ROWS.json",
296
+ "sha256": "6da5b5692a8a0cea3dc8f6d02d911fafd9bbf62e286a9be8db27c89f77a897e4"
297
  },
298
  "evidence": [
299
  {
300
+ "path": "results/accelerated-transfer-baselines-v1/decider/worker/COMPLETE.json",
301
+ "sha256": "e1ee0df8df59ec3eb2b49ba947a8508994a24bce226b7f2cccbe25f6835b99af"
302
  },
303
  {
304
+ "path": "results/accelerated-transfer-baselines-v1/decider/worker/METADATA.json",
305
+ "sha256": "f00061c1b0bb642bf1353aae867d19df988c01b0428b58206b31249521bcbff8"
306
  },
307
  {
308
+ "path": "results/accelerated-transfer-baselines-v1/decider/worker/predictions.jsonl",
309
+ "sha256": "f69e87fa65b09289e6861c1d789270aea5a92d23ac47817342647ed08f406df7"
310
  },
311
  {
312
+ "path": "analysis/decision-benchmark-v3-baselines/decider/REPORT.json",
313
+ "sha256": "5aae029f34524dfebc8e79128da8c39229c5747461dfce07ca59956da4416951"
314
  },
315
  {
316
+ "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
317
+ "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
318
+ },
319
+ {
320
+ "path": "analysis/decision-benchmark-v3-baselines/decider/COMPOSABLE-MANIFEST.json",
321
+ "sha256": "a49848e86feafaa4f36a33287910035e7605e344791aae139305b78cb304a77f"
322
  }
323
  ]
324
  },
325
+ "Qwen3.5-4B": {
326
  "panels": {
327
  "old_core": {
328
+ "path": "eval/heldout/primary-v2/base-4b/core/normalized.jsonl",
329
+ "sha256": "d20ad845066ff81caaf60ebee44df941872bdc642da78b5f3b3da4a33dd9a08d"
330
  },
331
  "v3_core": {
332
+ "path": "eval/heldout/v3/baselines/base-4b/core/normalized.jsonl",
333
+ "sha256": "7541893d2b2ec0ec2ffc513f50c568eaef860b65bcef9a1747996c116fff9661"
334
  },
335
  "v4": {
336
+ "path": "eval/heldout/v4/baselines/base-4b/core/normalized.jsonl",
337
+ "sha256": "7a6d252ccaa72bf54b94445190bbad023fb465a3eaf4cf657532e2cdd4f750f5"
338
  },
339
  "v5": {
340
+ "path": "eval/heldout/v5/open-baselines-v1/base-4b/core/normalized.jsonl",
341
+ "sha256": "c01ad5daed65267e9aec40ed9e97ece1b6511e4cab17ebff09339b41524fb7f1"
342
  }
343
  },
344
  "transfer_rows": {
345
+ "path": "analysis/decision-benchmark-v3-baselines/base-4b/ROWS.json",
346
+ "sha256": "37ab17d94d640ee46283e1f23723b9ea05133c22a8a9b62a97c756a89c5de785"
347
  },
348
  "evidence": [
349
  {
350
+ "path": "results/remaining-transfer-baselines-v1/base-4b/worker/COMPLETE.json",
351
+ "sha256": "72abc4f87a9234f95eb3b0a72a9e7d66798779d5982c757faffb610003b1b242"
352
  },
353
  {
354
+ "path": "results/remaining-transfer-baselines-v1/base-4b/worker/METADATA.json",
355
+ "sha256": "871fed92cf762187015e59174a042808dbf3fc88f7791aa0ddd530ae923196ab"
356
  },
357
  {
358
+ "path": "results/remaining-transfer-baselines-v1/base-4b/worker/predictions.jsonl",
359
+ "sha256": "8c47edda469f8914eccc2ecafa91c245752812e830cb9f29445653102522854a"
360
  },
361
  {
362
+ "path": "analysis/decision-benchmark-v3-baselines/base-4b/REPORT.json",
363
+ "sha256": "c2a0aafbe2b7c85b3c07a4b9843182a3361654f57883311155aa3b05cfc7002d"
364
  },
365
  {
366
  "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
367
  "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
368
  },
369
  {
370
+ "path": "analysis/decision-benchmark-v3-baselines/base-4b/COMPOSABLE-MANIFEST.json",
371
+ "sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
372
+ }
373
+ ]
374
+ },
375
+ "Sol": {
376
+ "panels": {
377
+ "old_core": {
378
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/old_core/normalized.jsonl",
379
+ "sha256": "5fc0b3e49a5de576f7621839130855c178b12d942dc357aaf6b4dd9735e0a6a8"
380
+ },
381
+ "v3_core": {
382
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v3_core/normalized.jsonl",
383
+ "sha256": "b2c99fbd739b849a66a55929a658de4517a7ffaa4496255566e6972ba28d4f15"
384
+ },
385
+ "v4": {
386
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v4/normalized.jsonl",
387
+ "sha256": "3c28302750e02d4cead0aacd66a323734e773f726c4c41db658dd6461251b28b"
388
+ },
389
+ "v5": {
390
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v5/normalized.jsonl",
391
+ "sha256": "6ddb0496012b596ea122afa4db4c4158e57dfb934e9ae29eb006756b652e31e1"
392
+ }
393
+ },
394
+ "transfer_rows": {
395
+ "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/v9-transfer-v9-test-published-temperature-rows.json",
396
+ "sha256": "b05799417cb3715de633135c58f713960b75bdb9ed30eb8912a197d33ee3d85b"
397
+ },
398
+ "evidence": [
399
+ {
400
+ "path": "results/expanded/sol-composition-v2/QUALITY-COLLECTION.json",
401
+ "sha256": "9b77aac87b85cbf3a74e86af4a2b9a03fd9e1ed854356a92ff54ccd0d47a16f4"
402
+ },
403
+ {
404
+ "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/REPORT.json",
405
+ "sha256": "5d8ebb339293687cec9089f54fb96aabf3330f16111487a49489f5e52bbb6f29"
406
+ }
407
+ ]
408
+ },
409
+ "kev-0.8b": {
410
+ "panels": {
411
+ "old_core": {
412
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-old_core-normalized.jsonl",
413
+ "sha256": "b492b4512ac164f83a319f31bbc9e086cae12f87ffe2c14f307481ffb240807c"
414
+ },
415
+ "v3_core": {
416
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v3_core-normalized.jsonl",
417
+ "sha256": "70cbfc69dfc72ac0c5ae01305ec63a9f31129c115b7068c6824cd22ead013d68"
418
+ },
419
+ "v4": {
420
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v4-normalized.jsonl",
421
+ "sha256": "761dfd69acbface101f2c21fd27d2db684c5c8b8cc425890a075729ae1f7db24"
422
+ },
423
+ "v5": {
424
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v5-normalized.jsonl",
425
+ "sha256": "51b6fc5f826ca8db44df43885f34d857ddd5b5a66453bbdc1a83fcf488928bdb"
426
+ }
427
+ },
428
+ "transfer_rows": {
429
+ "path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/v9-transfer-v9-test-published-temperature-rows.json",
430
+ "sha256": "d6abed88001aeac630219466f02c1aad89fcbd354b3ae533507f3c2f7fe42898"
431
+ },
432
+ "evidence": [
433
+ {
434
+ "path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/REPORT.json",
435
+ "sha256": "6760dcbaa7d85e6c2e893e4c80ab8e66e0b84d7cfa26ae9e25909f48a05f8de4"
436
+ },
437
+ {
438
+ "path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
439
+ "sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
440
  }
441
  ]
442
  },
 
490
  }
491
  ]
492
  },
493
+ "Laya-base": {
494
  "panels": {
 
 
 
 
 
 
 
 
495
  "v4": {
496
+ "path": "results/public-roster-baselines-v1/laya-english/worker/v4/normalized.jsonl",
497
+ "sha256": "31da1d4140da69d2677172707eb301843cf8fdfb43200cba93be910321f922ba"
498
  },
499
  "v5": {
500
+ "path": "results/public-roster-baselines-v1/laya-english/worker/v5/normalized.jsonl",
501
+ "sha256": "159614ceede814b578e1e025836a595c8c71a9bec4896caf6454d7bf54c4c44d"
502
+ },
503
+ "old_core": {
504
+ "path": "eval/heldout/primary-v2/laya-english/core/normalized.jsonl",
505
+ "sha256": "850643d2daeb47a6909985003736cab0f1f0e5a87631a35e27192f0733efd756"
506
+ },
507
+ "v3_core": {
508
+ "path": "eval/heldout/v3/baselines/laya-english/core/normalized.jsonl",
509
+ "sha256": "dec19caeb8074e25289e43a367392172da21f0c2ccdb403a059e2e4a2364a64b"
510
  }
511
  },
512
  "transfer_rows": {
513
+ "path": "analysis/decision-benchmark-v3-baselines/laya-english/ROWS.json",
514
+ "sha256": "19138909259400facb5d3d5402d0d736f870695fe3ee46019f2e438088c57275"
515
  },
516
  "evidence": [
517
  {
518
+ "path": "results/public-roster-baselines-v1/laya-english/worker/COMPLETE.json",
519
+ "sha256": "cef0da797ae14cebaf1e4b9edf9b16c3cd9a412a30435d7d2517b8ba6dbe8b98"
520
  },
521
  {
522
+ "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
523
+ "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
524
  },
525
  {
526
+ "path": "analysis/decision-benchmark-v3-baselines/laya-english/REPORT.json",
527
+ "sha256": "cb5935bb399dcf23ffe24f855452b2e074bcbb3d023cfc7dd439342e54f02c6b"
528
  },
529
  {
530
+ "path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
531
+ "sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
532
+ }
533
+ ]
534
+ },
535
+ "Laya-multilingual": {
536
+ "panels": {
537
+ "v4": {
538
+ "path": "results/public-roster-baselines-v1/laya-multilingual/worker/v4/normalized.jsonl",
539
+ "sha256": "814c1aa1f4d400942984ecb3786d3fd5fd15286d9b767a519c91993577795290"
540
+ },
541
+ "v5": {
542
+ "path": "results/public-roster-baselines-v1/laya-multilingual/worker/v5/normalized.jsonl",
543
+ "sha256": "6d4dbc16c7b977b15895e0e313dabc6d9db43df75ec56e89373d91f4f13ec676"
544
+ },
545
+ "old_core": {
546
+ "path": "eval/heldout/primary-v2/laya-multilingual/core/normalized.jsonl",
547
+ "sha256": "9ea39841c4d5901cd9ab860c980d2f3b55358be289c7e01b9b47982927bf849d"
548
  },
549
+ "v3_core": {
550
+ "path": "eval/heldout/v3/baselines/laya-multilingual/core/normalized.jsonl",
551
+ "sha256": "a0b7ce9daf4be2c7831bcab6b5e96c8dcb6bef133d39475f72241ab18faccdd2"
552
+ }
553
+ },
554
+ "transfer_rows": {
555
+ "path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/ROWS.json",
556
+ "sha256": "d32db6cf6576f287210d31336456497f9d6926179366db29304659baf9783b43"
557
+ },
558
+ "evidence": [
559
  {
560
+ "path": "results/public-roster-baselines-v1/laya-multilingual/worker/COMPLETE.json",
561
+ "sha256": "b190d683fc4c4762c2805a03d52628a53e125d228a3105c8ac2ddda00a5364e4"
562
  },
563
  {
564
+ "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
565
+ "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
566
+ },
567
+ {
568
+ "path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/REPORT.json",
569
+ "sha256": "d20d3d58f76ca2758545f83d358abd747f86d581c897d1fb521371d1000d6154"
570
+ },
571
+ {
572
+ "path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
573
+ "sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
574
  }
575
  ]
576
  }
577
  },
578
  "task_rows": 54,
579
  "rank_public_models": [
 
580
  "Lux",
581
  "Nox",
582
  "kev-9b",
 
588
  "kev-0.8b",
589
  "Qwen3.5-2B",
590
  "Laya-base",
591
+ "Laya-multilingual",
592
+ "Jev"
593
  ],
594
  "internal_extra_models_not_public_rank": [
595
  "llm2jev-2b",
 
672
  }
673
  },
674
  "source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
675
+ }
 
 
676
  }
model-card-example.json CHANGED
@@ -65,7 +65,7 @@
65
  "direct_engine_exact_response": true,
66
  "overflow_rejected": true,
67
  "overflow_message": "check: 40083 tokens exceeds max_length=16384; no truncation allowed",
68
- "bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
69
  "runtime": {
70
  "actual": {
71
  "torch": "2.12.0+git6bbd260",
 
65
  "direct_engine_exact_response": true,
66
  "overflow_rejected": true,
67
  "overflow_message": "check: 40083 tokens exceeds max_length=16384; no truncation allowed",
68
+ "bundle_manifest_sha256": "7d9b06bc25a75b1f131aabd280df0ef2777bf69db98574a300067ace40c61c9c",
69
  "runtime": {
70
  "actual": {
71
  "torch": "2.12.0+git6bbd260",
release-manifest.json CHANGED
@@ -1,12 +1,12 @@
1
  {
2
  "format": "decision-public-release-v1",
3
- "status": "documentation-only-assembled",
4
- "bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
5
  "readiness_sha256": "fcbb84a9bf7294ac79e2421cdd2cbfdfc46a79fd74c41ba0965778472d1cafc1",
6
- "model_card_sha256": "02678d253920bc54d7af864edbee273bb649a5268ede892a6602d6b023f66800",
7
  "repo_id": "llm-semantic-router/Decision-1.0-Nox",
8
- "assembly_script_sha256": "1697c4abaee22bb7b60259c042488178fc71847afe4810037683f874d7b8d513",
9
- "original_bundle_manifest_preserved": true,
10
  "files_exclude_this_manifest": true,
11
  "files": [
12
  {
@@ -17,7 +17,7 @@
17
  {
18
  "file": "DIAGNOSTICS.md",
19
  "bytes": 5678,
20
- "sha256": "f0e3a5040f9843bf1480f0383d2f01f7e9a601347d6b1545d68bf861d1f64007"
21
  },
22
  {
23
  "file": "Dockerfile.runtime",
@@ -26,8 +26,8 @@
26
  },
27
  {
28
  "file": "EVALUATION.md",
29
- "bytes": 3622,
30
- "sha256": "0624b6aebb35e86e8456ed734027ae4590fbf24c91775e4e228fe89454780f14"
31
  },
32
  {
33
  "file": "LICENSE",
@@ -36,14 +36,19 @@
36
  },
37
  {
38
  "file": "MATERIALS.json",
39
- "bytes": 2379,
40
- "sha256": "d52c25221dab8aa6ef1a3e8c91d6651a88c662246a00db4766f99af34e5a5271"
41
  },
42
  {
43
  "file": "NORMALIZATION_RUNTIME.md",
44
  "bytes": 756,
45
  "sha256": "cf8e6ce1f07687a68b6adeb98e6704b8f292cb6adc7b10e4b84e8e70aa5472e5"
46
  },
 
 
 
 
 
47
  {
48
  "file": "QUESTION-SCALING.md",
49
  "bytes": 1274,
@@ -56,13 +61,13 @@
56
  },
57
  {
58
  "file": "README.md",
59
- "bytes": 4672,
60
- "sha256": "02678d253920bc54d7af864edbee273bb649a5268ede892a6602d6b023f66800"
61
  },
62
  {
63
  "file": "RUNTIME-RELEASE.json",
64
- "bytes": 1264,
65
- "sha256": "97c0c58b97685f95ad57a65b378f4e4f374ff3f7a7e3f7ef37c7960131981ffe"
66
  },
67
  {
68
  "file": "RUNTIME.md",
@@ -76,8 +81,8 @@
76
  },
77
  {
78
  "file": "SENSITIVITY.md",
79
- "bytes": 751,
80
- "sha256": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048"
81
  },
82
  {
83
  "file": "SERVING_OPTIMIZATION.json",
@@ -91,13 +96,18 @@
91
  },
92
  {
93
  "file": "TASKS.md",
94
- "bytes": 9449,
95
- "sha256": "ffd8e2da2bed770855e69d54c44479bd66bf27a319684954ad8b55f742706341"
96
  },
97
  {
98
  "file": "USAGE.md",
99
- "bytes": 5435,
100
- "sha256": "46ec1078d2bb8cb54a1850809a101f07a77bb54c18b92276cb89d40245dbbc8f"
 
 
 
 
 
101
  },
102
  {
103
  "file": "assets/architecture.png",
@@ -236,18 +246,18 @@
236
  },
237
  {
238
  "file": "assets/decision-matrix.pdf",
239
- "bytes": 28175,
240
- "sha256": "12f963382e810eb2adb8a9fbbc3b3b7e25e75e4c13d56a04db81419e0b565033"
241
  },
242
  {
243
  "file": "assets/decision-matrix.png",
244
- "bytes": 324709,
245
- "sha256": "29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4"
246
  },
247
  {
248
  "file": "assets/decision-matrix.svg",
249
  "bytes": 43846,
250
- "sha256": "95e90ed75b9d942e212b621a2968be3c5387d0bf637ca681c885062c5de93456"
251
  },
252
  {
253
  "file": "assets/decision-question-scaling-600px.png",
@@ -271,18 +281,18 @@
271
  },
272
  {
273
  "file": "assets/decision-ranking.pdf",
274
- "bytes": 24211,
275
- "sha256": "2c5d77b3a2420531e25a35e74389cd74968840117f90ab9f0f79e3e245101a97"
276
  },
277
  {
278
  "file": "assets/decision-ranking.png",
279
- "bytes": 214771,
280
- "sha256": "a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f"
281
  },
282
  {
283
  "file": "assets/decision-ranking.svg",
284
  "bytes": 14180,
285
- "sha256": "39be265fee1f5e6a0b051524fadc33786c0a923a28d771a774ea3c01ace460dc"
286
  },
287
  {
288
  "file": "assets/readout.png",
@@ -316,8 +326,8 @@
316
  },
317
  {
318
  "file": "bundle-manifest.json",
319
- "bytes": 7992,
320
- "sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830"
321
  },
322
  {
323
  "file": "chat_template.jinja",
@@ -326,8 +336,8 @@
326
  },
327
  {
328
  "file": "code/decision_api.py",
329
- "bytes": 10952,
330
- "sha256": "273f6f10f22d5a68b8db34cfcbd35407fb43d8030f6d7cd188bcf118dd90a152"
331
  },
332
  {
333
  "file": "code/decision_model.py",
@@ -361,8 +371,8 @@
361
  },
362
  {
363
  "file": "metrics/benchmark.json",
364
- "bytes": 2276797,
365
- "sha256": "40b1da632cb586a5122fa34f85ea654e2890d6619bbe707b49afe829b4e8d3ce"
366
  },
367
  {
368
  "file": "metrics/comparator-coverage.json",
@@ -371,8 +381,8 @@
371
  },
372
  {
373
  "file": "metrics/evaluation-provenance.json",
374
- "bytes": 28685,
375
- "sha256": "cab27fc814933826d4a280ab280877c4683524a5ea67773e4e134a00fcf8befa"
376
  },
377
  {
378
  "file": "metrics/expanded-quality.json",
@@ -407,7 +417,7 @@
407
  {
408
  "file": "model-card-example.json",
409
  "bytes": 3775,
410
- "sha256": "6bb34ba80660291be6c53005571a8a1d6546852bca186fff5da915019195d3a9"
411
  },
412
  {
413
  "file": "pyproject.toml",
@@ -557,10 +567,10 @@
557
  "MATERIALS.json": "copy",
558
  "metrics/semantic-consistency.json": "copy"
559
  },
560
- "scope": "13model public display only; original benchmark values and all model files unchanged.",
561
- "release_tag": "v1.3.1",
562
- "change_kind": "documentation-only-public-roster",
563
- "previous_main_revision": "8aea799ad37bf78ea8d3be42bde29c324524889e",
564
- "previous_release_manifest_sha256": "752ff577b903fb71e1012d1b0a9ccafc83336e0a104126a7ff918389c3673198",
565
  "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
566
  }
 
1
  {
2
  "format": "decision-public-release-v1",
3
+ "status": "qualified-runtime-and-current-documents-assembled",
4
+ "bundle_manifest_sha256": "7d9b06bc25a75b1f131aabd280df0ef2777bf69db98574a300067ace40c61c9c",
5
  "readiness_sha256": "fcbb84a9bf7294ac79e2421cdd2cbfdfc46a79fd74c41ba0965778472d1cafc1",
6
+ "model_card_sha256": "41dfa047d33bdc3aa9feff55928972533ba803a30b332ca380f042268fa1e848",
7
  "repo_id": "llm-semantic-router/Decision-1.0-Nox",
8
+ "assembly_script_sha256": "a8f72b8a67a64034d6183f89c945bab611056816ba0cba43670e9becb6eb78c4",
9
+ "original_bundle_manifest_preserved": false,
10
  "files_exclude_this_manifest": true,
11
  "files": [
12
  {
 
17
  {
18
  "file": "DIAGNOSTICS.md",
19
  "bytes": 5678,
20
+ "sha256": "0b290793565cc5007e9ee7bc79625812c73202c7d8bfce79d8816a0e7a93a965"
21
  },
22
  {
23
  "file": "Dockerfile.runtime",
 
26
  },
27
  {
28
  "file": "EVALUATION.md",
29
+ "bytes": 3705,
30
+ "sha256": "4159b7939a107e91b13e47345a192eea489f8bd870943e4b397e094b33bba905"
31
  },
32
  {
33
  "file": "LICENSE",
 
36
  },
37
  {
38
  "file": "MATERIALS.json",
39
+ "bytes": 2111,
40
+ "sha256": "6591d0318677c51084db78bd3bfc09bb5cacaa6c9c2d5700fc969f64b9a7ac4c"
41
  },
42
  {
43
  "file": "NORMALIZATION_RUNTIME.md",
44
  "bytes": 756,
45
  "sha256": "cf8e6ce1f07687a68b6adeb98e6704b8f292cb6adc7b10e4b84e8e70aa5472e5"
46
  },
47
+ {
48
+ "file": "NULL_DESCRIPTION_RENDERING.json",
49
+ "bytes": 514,
50
+ "sha256": "3b531cab60cba35648ab71fe4be7fd1af619f02ee337e3beed6296fdfc8b9ec8"
51
+ },
52
  {
53
  "file": "QUESTION-SCALING.md",
54
  "bytes": 1274,
 
61
  },
62
  {
63
  "file": "README.md",
64
+ "bytes": 4767,
65
+ "sha256": "41dfa047d33bdc3aa9feff55928972533ba803a30b332ca380f042268fa1e848"
66
  },
67
  {
68
  "file": "RUNTIME-RELEASE.json",
69
+ "bytes": 679,
70
+ "sha256": "325d7e037acab5f61dab4d56d817e7d7e2e873d9b5f0d0df587afd6c33364a46"
71
  },
72
  {
73
  "file": "RUNTIME.md",
 
81
  },
82
  {
83
  "file": "SENSITIVITY.md",
84
+ "bytes": 788,
85
+ "sha256": "c3ea31420db4f2e5f17cc7197197a3ddd18abdad773ffb5bf1171b67136c3bc2"
86
  },
87
  {
88
  "file": "SERVING_OPTIMIZATION.json",
 
96
  },
97
  {
98
  "file": "TASKS.md",
99
+ "bytes": 9299,
100
+ "sha256": "2fb04bc9c17d805e2abd5e90520df25e79d85aa5296cd42074d4b156f54bf862"
101
  },
102
  {
103
  "file": "USAGE.md",
104
+ "bytes": 5665,
105
+ "sha256": "b4c2f44ce120149c7cdee3749283958f34ba9caf4b6492442818c1b6a748c0d3"
106
+ },
107
+ {
108
+ "file": "WEIGHTING.md",
109
+ "bytes": 788,
110
+ "sha256": "c3ea31420db4f2e5f17cc7197197a3ddd18abdad773ffb5bf1171b67136c3bc2"
111
  },
112
  {
113
  "file": "assets/architecture.png",
 
246
  },
247
  {
248
  "file": "assets/decision-matrix.pdf",
249
+ "bytes": 28045,
250
+ "sha256": "5e127c0cd1a6f83b772a3e66ac2e1f4ec19913ac4938371c6f190a23c50a29fe"
251
  },
252
  {
253
  "file": "assets/decision-matrix.png",
254
+ "bytes": 325781,
255
+ "sha256": "c775737aedeecadb0824092d37014a8e2cffa90af0dafcf10d584e0db823ecf8"
256
  },
257
  {
258
  "file": "assets/decision-matrix.svg",
259
  "bytes": 43846,
260
+ "sha256": "d99102fafe842f2a6bdd4d455f915833d5e0e09e3a22e5211fbf86e5647247ab"
261
  },
262
  {
263
  "file": "assets/decision-question-scaling-600px.png",
 
281
  },
282
  {
283
  "file": "assets/decision-ranking.pdf",
284
+ "bytes": 24204,
285
+ "sha256": "a28d4b0e8649cb7bf2db452a11d93675631af795561726aa63432b05b5888049"
286
  },
287
  {
288
  "file": "assets/decision-ranking.png",
289
+ "bytes": 215215,
290
+ "sha256": "d252593501e8faeb9c5c8cc0a9ae9a96591e7aa8696cf5394cdfec9c0ddac3b7"
291
  },
292
  {
293
  "file": "assets/decision-ranking.svg",
294
  "bytes": 14180,
295
+ "sha256": "ab7ca66a9202587dee593714e46be760cb12ac91b6459074bcbbe11c59107d9d"
296
  },
297
  {
298
  "file": "assets/readout.png",
 
326
  },
327
  {
328
  "file": "bundle-manifest.json",
329
+ "bytes": 8171,
330
+ "sha256": "7d9b06bc25a75b1f131aabd280df0ef2777bf69db98574a300067ace40c61c9c"
331
  },
332
  {
333
  "file": "chat_template.jinja",
 
336
  },
337
  {
338
  "file": "code/decision_api.py",
339
+ "bytes": 10978,
340
+ "sha256": "6716bdabca3cd2f1aa447d42a6cf53ee62d7ea97984a01f16b4c09fac8aaf17e"
341
  },
342
  {
343
  "file": "code/decision_model.py",
 
371
  },
372
  {
373
  "file": "metrics/benchmark.json",
374
+ "bytes": 2261123,
375
+ "sha256": "178e2e5f58da45f9fbf378ca82b49d7878fc8e9c93eee195a5ef1e7ba0f4fc6b"
376
  },
377
  {
378
  "file": "metrics/comparator-coverage.json",
 
381
  },
382
  {
383
  "file": "metrics/evaluation-provenance.json",
384
+ "bytes": 29254,
385
+ "sha256": "a3718fcee42a8cfe0144b01e04346c36a09ac7958154e4561327753d87d5fe8c"
386
  },
387
  {
388
  "file": "metrics/expanded-quality.json",
 
417
  {
418
  "file": "model-card-example.json",
419
  "bytes": 3775,
420
+ "sha256": "1c4fc87d543722e2c3b839cbebd30c5d1370fcaf2d897f97de13079a668ac93e"
421
  },
422
  {
423
  "file": "pyproject.toml",
 
567
  "MATERIALS.json": "copy",
568
  "metrics/semantic-consistency.json": "copy"
569
  },
570
+ "scope": "Latest13model five-panel comparison and54task diagnostics; model tensor/tokenizer/temperature identities unchanged.",
571
+ "release_tag": "v1.3.2",
572
+ "change_kind": "null-Choice-description-runtime-and-current-materials",
573
+ "previous_main_revision": "46505c737a45cbe2c4ac4eee38e4f94cb520e4ed",
574
+ "previous_release_manifest_sha256": "f537b023f52d079fdd77ecb5b1240799959209de172cc6162a6996a009ef7d03",
575
  "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
576
  }