Xunzhuo commited on
Commit
8aea799
·
verified ·
1 Parent(s): 74979f7

Publish complete decision benchmark, task diagnostics, and current latency

Browse files
.gitattributes CHANGED
@@ -53,3 +53,5 @@ assets/decision-expanded-v3_core.png filter=lfs diff=lfs merge=lfs -text
53
  assets/decision-expanded-v4.png filter=lfs diff=lfs merge=lfs -text
54
  assets/decision-expanded-v5.png filter=lfs diff=lfs merge=lfs -text
55
  assets/decision-question-scaling.png filter=lfs diff=lfs merge=lfs -text
 
 
 
53
  assets/decision-expanded-v4.png filter=lfs diff=lfs merge=lfs -text
54
  assets/decision-expanded-v5.png filter=lfs diff=lfs merge=lfs -text
55
  assets/decision-question-scaling.png filter=lfs diff=lfs merge=lfs -text
56
+ assets/decision-matrix.png filter=lfs diff=lfs merge=lfs -text
57
+ assets/decision-ranking.png filter=lfs diff=lfs merge=lfs -text
DIAGNOSTICS.md ADDED
@@ -0,0 +1,108 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Diagnostics
2
+
3
+ These axes remain separate from headline accuracy. Probability metrics use the shipped distribution; no temperature is fitted on this benchmark. An unavailable full-population metric is shown as —, without replacing it with a successful-only average.
4
+
5
+ ## Probability quality on transfer
6
+
7
+ | Model | Valid / requested | Brier ↓ | NLL ↓ | ECE % ↓ | Coverage at ≤5% error % ↑ | AURC ↓ |
8
+ |---|---:|---:|---:|---:|---:|---:|
9
+ | Jev | 1046/1046 | 0.1912 | 0.5743 | 3.35 | 76.96 | 0.0346 |
10
+ | Lux | 1046/1046 | 0.3058 | 0.6201 | 5.13 | 57.36 | 0.0639 |
11
+ | Nox | 1046/1046 | 0.4350 | 0.9375 | 10.27 | 35.18 | 0.1195 |
12
+ | Kev-9B | 1046/1046 | 0.2948 | 0.6028 | 4.45 | 58.03 | 0.0586 |
13
+ | Kev-4B | 1046/1046 | 0.3145 | 0.6434 | 2.61 | 56.79 | 0.0659 |
14
+ | Qwen3.5-9B | 1046/1046 | 0.3673 | 0.7597 | 9.75 | 47.42 | 0.0872 |
15
+ | Decider | 1046/1046 | 0.4122 | 0.7969 | 8.09 | 35.18 | 0.1172 |
16
+ | Qwen3.5-4B | 1046/1046 | 0.4199 | 0.8186 | 8.98 | 34.70 | 0.1233 |
17
+ | Sol | 1046/1046 | 0.5454 | 1.1094 | 11.25 | 16.16 | 0.2073 |
18
+ | Kev-0.8B | 1046/1046 | 0.4839 | 0.9489 | 2.58 | 21.61 | 0.1782 |
19
+ | Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
20
+ | Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
21
+ | Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
22
+ | Kai | 1046/1046 | 0.6066 | 1.1552 | 7.95 | 1.24 | 0.3343 |
23
+
24
+ Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
25
+
26
+ ## Option-order robustness
27
+
28
+ | Model | Valid / requested pairs | Both correct % ↑ | Semantic flip % ↓ | Mean half-L1 ↓ |
29
+ |---|---:|---:|---:|---:|
30
+ | Jev | 36/36 | 86.11 | 0.00 | 0.0208 |
31
+ | Lux | 36/36 | 77.78 | 8.33 | 0.0923 |
32
+ | Nox | 36/36 | 58.33 | 25.00 | 0.1049 |
33
+ | Kev-9B | 36/36 | 80.56 | 2.78 | 0.0615 |
34
+ | Kev-4B | 36/36 | 77.78 | 5.56 | 0.0667 |
35
+ | Qwen3.5-9B | 36/36 | 75.00 | 11.11 | 0.1157 |
36
+ | Decider | 36/36 | 83.33 | 11.11 | 0.0701 |
37
+ | Qwen3.5-4B | 36/36 | 77.78 | 13.89 | 0.1390 |
38
+ | Sol | 36/36 | 41.67 | 38.89 | 0.0694 |
39
+ | Kev-0.8B | 36/36 | 55.56 | 16.67 | 0.0808 |
40
+ | Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
41
+ | Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
42
+ | Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
43
+ | Kai | 36/36 | 41.67 | 27.78 | 0.1170 |
44
+
45
+ The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
46
+
47
+ ## Missing evidence
48
+
49
+ | Model | Valid / requested | Intact/control accuracy % ↑ | Mean max P % ↓ | P≥0.9 share % ↓ | Normalized entropy ↑ | Paired confidence drop pp ↑ |
50
+ |---|---:|---:|---:|---:|---:|---:|
51
+ | Jev | 110/110 | 93.64 | 62.04 | 16.36 | 0.7657 | 30.30 |
52
+ | Lux | 110/110 | 84.55 | 68.81 | 19.09 | 0.6785 | 19.59 |
53
+ | Nox | 110/110 | 72.73 | 78.65 | 27.27 | 0.5072 | 12.64 |
54
+ | Kev-9B | 110/110 | 91.82 | 39.61 | 0.00 | 0.9981 | 53.19 |
55
+ | Kev-4B | 110/110 | 91.82 | 41.47 | 0.00 | 0.9918 | 51.57 |
56
+ | Qwen3.5-9B | 110/110 | 80.91 | 67.62 | 7.27 | 0.7585 | 19.09 |
57
+ | Decider | 110/110 | 69.09 | 68.82 | 16.36 | 0.6772 | 13.59 |
58
+ | Qwen3.5-4B | 110/110 | 72.73 | 59.89 | 0.91 | 0.8120 | 22.35 |
59
+ | Sol | 110/110 | 65.45 | 77.19 | 24.55 | 0.5424 | 5.68 |
60
+ | Kev-0.8B | 110/110 | 87.27 | 42.11 | 0.00 | 0.9796 | 40.34 |
61
+ | Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
62
+ | Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
63
+ | Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
64
+ | Kai | 110/110 | 37.27 | 55.39 | 18.18 | 0.8262 | -0.07 |
65
+
66
+ The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
67
+
68
+ ## Native contract coverage
69
+
70
+ | Model | Original probability rows | Transfer probability rows | Transfer truncated questions |
71
+ |---|---:|---:|---:|
72
+ | Jev | 2720/2720 | 1264/1264 | — / not observable |
73
+ | Lux | 2720/2720 | 1264/1264 | 0 |
74
+ | Nox | 2720/2720 | 1264/1264 | 0 |
75
+ | Kev-9B | 2720/2720 | 1264/1264 | 0 |
76
+ | Kev-4B | 2720/2720 | 1264/1264 | 0 |
77
+ | Qwen3.5-9B | 2720/2720 | 1264/1264 | 0 |
78
+ | Decider | 2720/2720 | 1264/1264 | 0 |
79
+ | Qwen3.5-4B | 2720/2720 | 1264/1264 | 0 |
80
+ | Sol | 2720/2720 | 1264/1264 | 0 |
81
+ | Kev-0.8B | 2720/2720 | 1264/1264 | 0 |
82
+ | Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
83
+ | Laya · English | 2720/2720 | 1264/1264 | 34 |
84
+ | Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
85
+ | Kai | 2720/2720 | 1264/1264 | 0 |
86
+
87
+ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation; Kai retains the shipped complete-request 1,024-token limit. Server-side truncation for Jev cannot be observed.
88
+
89
+ ## Uncertainty
90
+
91
+ | Model | Overall % | 95% component-bootstrap interval |
92
+ |---|---:|---:|
93
+ | Jev | 81.05 | 79.70–82.35 |
94
+ | Lux | 76.72 | 75.35–78.07 |
95
+ | Nox | 72.84 | 71.33–74.31 |
96
+ | Kev-9B | 71.89 | 70.42–73.35 |
97
+ | Kev-4B | 70.09 | 68.45–71.63 |
98
+ | Qwen3.5-9B | 69.73 | 68.27–71.20 |
99
+ | Decider | 67.71 | 66.10–69.34 |
100
+ | Qwen3.5-4B | 67.29 | 65.89–68.69 |
101
+ | Sol | 66.32 | 64.77–67.85 |
102
+ | Kev-0.8B | 58.28 | 56.64–59.89 |
103
+ | Qwen3.5-2B | 57.24 | 55.74–58.76 |
104
+ | Laya · English | 51.03 | 49.43–52.68 |
105
+ | Laya · Multilingual | 47.19 | 45.58–48.82 |
106
+ | Kai | 46.49 | 45.02–47.99 |
107
+
108
+ Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
EVALUATION.md CHANGED
@@ -1,124 +1,38 @@
1
  # Evaluation
2
 
3
- This page compares the latest published Decision releases with fixed external references. All displayed cells are copied from verified completed measurements; no predictions, confidence intervals, task weights or release decisions were recomputed for this documentation update.
4
 
5
- ## Quality protocol
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6
 
7
- The headline averages four panels equally: 880 general decisions, 880 compositional tasks, 480 natural-reading questions and 480 reading/inference questions. Original within-panel family/source weights are retained; this is not an unweighted average of all 27 tasks. Two native-interface supplements are outside the 2,720-decision mean. Confidence intervals use the original whole-component paired bootstrap, including shared language and counterfactual groups.
8
 
9
- The current model is shown first, the other current Decision model second, external open references next, and Jev last as a closed-service frontier reference. This display order is not a ranking. Bold uses unrounded values and the strict external-open-reference rule described in the card; exact ties are not bold.
10
 
11
- | Model | Version | General decisions | Compositional tasks | Natural reading | Reading and inference | Mean accuracy ↑ |
12
- |---|---|---:|---:|---:|---:|---:|
13
- | Nox | 4B · v1.3 | **83.00** | **51.79** | 79.06 | **86.25** | **75.03** |
14
- | Sol | 2B · v1.3 | 73.75 | 46.08 | 76.56 | 84.17 | 70.14 |
15
- | Kev | 9B | 76.36 | 45.54 | 86.72 | 83.75 | 73.09 |
16
- | Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 71.75 |
17
- | Qwen3.5 | 4B · untuned | 69.89 | 43.33 | 87.97 | 79.79 | 70.25 |
18
- | Qwen3.5 | 2B · untuned | 57.12 | 39.00 | 73.75 | 72.29 | 60.54 |
19
- | Laya | Upstream default | 57.01 | 37.75 | 51.25 | 63.75 | 52.44 |
20
- | Jev | 1.13.0 · frontier | 79.10 | 66.38 | 94.53 | 89.79 | 82.45 |
21
 
22
- Accuracy (%). Mean weights each panel equally; each panel retains its frozen family/source weights.
23
- Bold marks Sol or Nox strictly above every external open reference in that metric; Jev and the other Decision model are excluded from the threshold. Ties are not bold.
24
 
25
- ### General decisions
26
 
27
- | Task | Nox · 4B · v1.3 | Sol · 2B · v1.3 | Kev · 9B | Decider · 2B | Qwen3.5 · 4B · untuned | Qwen3.5 · 2B · untuned | Laya · Upstream default | Jev · 1.13.0 · frontier |
28
- |---|---:|---:|---:|---:|---:|---:|---:|---:|
29
- | News classification | 85.16 | 83.59 | 88.28 | 86.72 | 84.38 | 80.47 | 91.41 | 85.16 |
30
- | Boolean constraints | **93.75** | 50.00 | 92.19 | 71.88 | 62.50 | 37.50 | 43.75 | 100.00 |
31
- | Entity classification | 96.43 | 95.54 | 98.21 | 98.21 | 97.32 | 92.86 | 83.93 | 96.43 |
32
- | Intent routing | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 85.94 | 100.00 |
33
- | Evidence placement | **100.00** | **98.96** | 50.00 | 8.33 | 72.92 | 9.38 | 95.83 | 30.21 |
34
- | Ordered rubric | 89.06 | 65.62 | 96.88 | 84.38 | 90.62 | 81.25 | 18.75 | 100.00 |
35
- | Relation composition | 52.08 | 37.50 | 52.08 | 54.17 | 51.04 | 48.96 | 25.00 | 56.25 |
36
- | Scoped evidence | **78.12** | **79.17** | 63.54 | 47.92 | 51.04 | 36.46 | 37.50 | 89.58 |
37
- | State tracking | 35.42 | 27.08 | 36.46 | 29.17 | 28.12 | 25.00 | 23.96 | 33.33 |
38
- | In / out of menu | **100.00** | **100.00** | 85.94 | 59.38 | 60.94 | 62.50 | 64.06 | 100.00 |
39
 
40
- ### Compositional tasks
41
 
42
- | Task | Nox · 4B · v1.3 | Sol · 2B · v1.3 | Kev · 9B | Decider · 2B | Qwen3.5 · 4B · untuned | Qwen3.5 · 2B · untuned | Laya · Upstream default | Jev · 1.13.0 · frontier |
43
- |---|---:|---:|---:|---:|---:|---:|---:|---:|
44
- | Record identity | 53.75 | 50.00 | 46.25 | 56.25 | 50.00 | 46.25 | 47.50 | 76.25 |
45
- | Capacity assignment | 46.25 | 47.50 | 50.00 | 51.25 | 50.00 | 50.00 | 63.75 | 68.75 |
46
- | Constraint assignment | **41.25** | **26.25** | 25.00 | 23.75 | 23.75 | 21.25 | 22.50 | 55.00 |
47
- | Intent routing · EN | 91.67 | 90.83 | 86.67 | 90.00 | 91.67 | 70.83 | 74.17 | 91.67 |
48
- | Intent routing · ZH | 87.50 | 87.50 | 85.00 | 88.33 | 86.67 | 74.17 | 73.33 | 88.33 |
49
- | Multiset reconciliation | 28.75 | 28.75 | 35.00 | 23.75 | 27.50 | 27.50 | 26.25 | 58.75 |
50
- | Ordered service loss | **30.00** | **25.00** | 21.25 | 22.50 | 21.25 | 20.00 | 20.00 | 42.50 |
51
- | Conflicting rule closure | **33.75** | 25.00 | 28.75 | 26.25 | 25.00 | 27.50 | 18.75 | 66.25 |
52
- | Temporal exclusion | 42.50 | 41.25 | 32.50 | 42.50 | 31.25 | 28.75 | 12.50 | 41.25 |
53
- | Transaction recovery | **62.50** | 38.75 | 45.00 | 41.25 | 26.25 | 23.75 | 18.75 | 75.00 |
54
 
55
- ### Natural reading
56
-
57
- | Task | Nox · 4B · v1.3 | Sol · 2B · v1.3 | Kev · 9B | Decider · 2B | Qwen3.5 · 4B · untuned | Qwen3.5 · 2B · untuned | Laya · Upstream default | Jev · 1.13.0 · frontier |
58
- |---|---:|---:|---:|---:|---:|---:|---:|---:|
59
- | Yes / no reading | 86.25 | 83.12 | 90.62 | 91.88 | 83.75 | 65.00 | 69.38 | 92.50 |
60
- | Reading · EN | 73.12 | 70.00 | 83.75 | 93.12 | 93.12 | 83.75 | 38.12 | 96.88 |
61
- | Reading · ZH | 70.62 | 70.00 | 81.88 | 91.25 | 91.25 | 81.25 | 28.12 | 96.25 |
62
-
63
- ### Reading and inference
64
-
65
- | Task | Nox · 4B · v1.3 | Sol · 2B · v1.3 | Kev · 9B | Decider · 2B | Qwen3.5 · 4B · untuned | Qwen3.5 · 2B · untuned | Laya · Upstream default | Jev · 1.13.0 · frontier |
66
- |---|---:|---:|---:|---:|---:|---:|---:|---:|
67
- | Contextual reasoning | **70.83** | **69.17** | 67.50 | 66.67 | 60.83 | 56.67 | 30.00 | 86.67 |
68
- | Answerability | **88.33** | 83.33 | 80.83 | 84.17 | 81.67 | 82.50 | 57.50 | 90.83 |
69
- | Textual entailment | 89.17 | 88.33 | 90.00 | 90.83 | 80.00 | 64.17 | 72.50 | 82.50 |
70
- | Scientific inference | 96.67 | 95.83 | 96.67 | 95.83 | 96.67 | 85.83 | 95.00 | 99.17 |
71
-
72
- ### Probability quality
73
-
74
- | Model | Version | Weighted Brier ↓ | Valid probability rows |
75
- |---|---|---:|---:|
76
- | Nox | 4B · v1.3 | **0.3323** | 2,720 / 2,720 |
77
- | Sol | 2B · v1.3 | 0.4258 | 2,720 / 2,720 |
78
- | Kev | 9B | 0.3715 | 2,720 / 2,720 |
79
- | Decider | 2B | 0.3535 | 2,720 / 2,720 |
80
- | Qwen3.5 | 4B · untuned | 0.3996 | 2,720 / 2,720 |
81
- | Qwen3.5 | 2B · untuned | 0.5130 | 2,720 / 2,720 |
82
- | Laya | Upstream default | 0.5903 | 2,720 / 2,720 |
83
- | Jev | 1.13.0 · frontier | 0.2287 | 2,720 / 2,720 |
84
-
85
- Missing Brier means probability coverage was incomplete; no supported-only average is substituted.
86
- The two native contract supplements are reported separately and are outside this quality mean.
87
-
88
- ## Current probability quality
89
-
90
- Nox v1.3 raw T=1 weighted Brier is 0.349587; the shipped CAL-fitted Brier is 0.332289. Calibration uses the fixed development CAL split after weights are selected. It does not improve category accuracy by itself, and confidence is not a factuality guarantee.
91
-
92
- ## Latency
93
-
94
- Only this release is displayed in [the latency tables and curve](QUESTION-SCALING.md). Tokenization and inference are included; loading and network are excluded. Three fresh processes supply 30 measured requests per point on an otherwise idle AMD gfx942 GPU. These measurements do not imply service concurrency or a cross-model throughput ranking.
95
-
96
- ## Comparator identities and scope
97
-
98
- [Nox v1.3](https://huggingface.co/llm-semantic-router/Decision-1.0-Nox/tree/ad089ad3a5dc9a7a21e6d96db546bb53e2212654) and [Sol v1.3](https://huggingface.co/llm-semantic-router/Decision-1.0-Sol/tree/2412d9470d3aa125b346262dad80f6161847aa09) use their actually published bundles and completed four-panel evaluations.
99
-
100
- [Kev-9B](https://huggingface.co/jaredpalmer/kev-9b/tree/6281032426a9ca3a08374d137bf9dfd8afd78e8a) uses its unchanged native adapter at source revision `e0bcf50153f1bda4ca6a8be5e12cbd5f9ebbce1c` under the recorded matched-FLA runtime. Its four quality panels have complete 2,720-row coverage; its separate native panels each succeed on 217 of 220 requests, so complete quality coverage must not be read as universal native-interface support.
101
-
102
- Untuned Qwen references use the frozen chat/LM-head adapter. Laya retains its upstream default language router and native truncation; Decider retains its native adapter. Jev is the recorded 1.13.0 service snapshot; its server-side truncation is not observable. These previously observed evaluation suites are not all pristine holdouts. Supplemental finite-probe scores are outside this mean and do not alter the release rule.
103
-
104
- The architecture, weights, tokenizer, normalization profile, temperature and original qualification are unchanged by this documentation refresh. Historical paired-update evidence remains archived as machine-readable provenance; no old release is presented as a current comparison row.
105
-
106
- [Exact displayed statistics](metrics/expanded-quality.json) · [Comparator provenance](metrics/comparator-coverage.json) · [Figure evidence](metrics/figure-evidence.json)
107
-
108
- ## Semantic consistency
109
-
110
- | Model | Consistent / eligible groups | Consistency ↑ |
111
- |---|---:|---:|
112
- | Nox · 4B · v1.3 | 174/192 | **90.62** |
113
- | Sol · 2B · v1.3 | 167/192 | **86.98** |
114
- | Kev · 9B | 143/192 | 74.48 |
115
- | Decider · 2B | 126/192 | 65.62 |
116
- | Qwen3.5 · 4B · untuned | 98/192 | 51.04 |
117
- | Qwen3.5 · 2B · untuned | 94/192 | 48.96 |
118
- | Laya · Upstream default | 115/192 | 59.90 |
119
- | Jev · 1.13.0 · frontier | 158/192 | 82.29 |
120
-
121
- Mixed equivalent variants of language, keys and presentation; **consistency is not accuracy**. A consistently wrong answer still counts. This observed-core supplement weights 192 eligible semantic groups equally and is outside the four-panel mean. [Definitions and exact counts](metrics/semantic-consistency.json).
122
-
123
-
124
- Eligible groups require more than one variant, uniquely identifiable option semantics, and a shared semantic gold. All outputs must be valid and agree semantically. The 240 singleton news/entity groups are excluded. This is a mixed-variation measure, not an isolated option-order experiment.
 
1
  # Evaluation
2
 
3
+ The comparison covers **3,766 scored decisions across 54 tasks** and all 14 displayed models. All models use the same requested examples, candidate order and scoring rules. Each model keeps its native admission, truncation and probability semantics.
4
 
5
+ | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
6
+ |---|---:|---:|---:|---:|---:|---:|---:|
7
+ | Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
8
+ | Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
9
+ | Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** |
10
+ | Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
11
+ | Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
12
+ | Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
13
+ | Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 |
14
+ | Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 |
15
+ | Sol | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 |
16
+ | Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 |
17
+ | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
18
+ | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
19
+ | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
20
+ | Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
21
 
22
+ Accuracy (%). Rows are sorted by overall score. Bold marks Decision-family cells strictly above every external open or untuned reference in that metric; Jev and the other Decision models are excluded from that threshold.
23
 
24
+ ## Scope and weighting
25
 
26
+ The overall score weights **Decisions 30%, Composition 25%, Reading 15%, Inference 15%, Transfer 15%**. The first four panels contain 880, 880, 480 and 480 questions and retain their original family/source weights. Transfer is micro-accuracy over 1,046 clean knowable questions from the frozen upstream transfer test; 110 missing-evidence questions and 108 variants remain separate diagnostics. No latency, calibration error or consistency score is averaged into accuracy.
 
 
 
 
 
 
 
 
 
27
 
28
+ These are **outcome-informed product-priority weights, chosen after observing benchmark results**. The data are observed regression tests, not a fresh blind test. Reweighting is not a training improvement. [Weight sensitivity](SENSITIVITY.md) retains the prior weighting and original four-panel comparison for the same model weights. Training, checkpoint selection and calibration do not use these test labels.
 
29
 
30
+ ## Full results
31
 
32
+ [All 54 task rows](TASKS.md) preserve every original decision, composition, reading and inference task plus all 27 transfer tasks. [Diagnostics](DIAGNOSTICS.md) separately report probability quality, option-order sensitivity, missing evidence, native coverage and uncertainty. [Exact statistics](metrics/benchmark.json) include counts and confidence intervals.
 
 
 
 
 
 
 
 
 
 
 
33
 
34
+ ## Model and API scope
35
 
36
+ Decision models return typed Choice, Noul and Score answers in the SystemOne format. Encoder and decoder references are evaluated through their published native interfaces. Untuned Qwen models use the frozen letter-logit readout, with no generated-text parsing. Kev uses the matched BF16 backbone with FP32 head and shipped temperature; date-fact injection is disabled. This differs from the authors’ default FP32 reproduction. Laya English and multilingual are measured separately. Kai uses its shipped 1,024-token complete-request limit. Jev is a recorded hosted-service snapshot.
 
 
 
 
 
 
 
 
 
 
 
37
 
38
+ The tests assess decisions from the supplied state and fixed choices. They do not establish live fact retrieval, universally superior reasoning or cross-hardware speed rankings. [Immutable identities and evaluation provenance](metrics/evaluation-provenance.json).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
MATERIALS.json CHANGED
@@ -1,224 +1,36 @@
1
  {
2
- "status": "docs-only-review-candidate-not-uploaded",
3
- "generator_sha256": "2de88a794052e4cad6d18e0c0fb3efa7aaa2a7d20e321a851906dec04a0ac81d",
4
- "source_evidence_sha256": "8ef8f73315fc185dca5c887728c44ed517a97623cbe25710c2dc83480204d3a8",
5
- "files": [
6
- {
7
- "file": "EVALUATION.md",
8
- "bytes": 9573,
9
- "sha256": "1c7e88ca75fe85502440dc4202de91786e3f68b8533e7901f398c5ee5352c4b6"
10
- },
11
- {
12
- "file": "QUESTION-SCALING.md",
13
- "bytes": 1555,
14
- "sha256": "fd5355fb6b69575a8812fa30cf1fcee2065182923206ab7dbe34a217db1bda9e"
15
- },
16
- {
17
- "file": "README.md",
18
- "bytes": 5001,
19
- "sha256": "4ec3943c766b36eed886ac2713363dbcbc5b736b942c1e38f13cb1ab5f5758cd"
20
- },
21
- {
22
- "file": "assets/architecture.png",
23
- "bytes": 511058,
24
- "sha256": "8fcfa10b3941ee4aec1f5fba20a25e382a7d5de83074071c9ef68e627996de40"
25
- },
26
- {
27
- "file": "assets/architecture.svg",
28
- "bytes": 14718,
29
- "sha256": "215d9af6f94cc24847fc1d23a0b2a287b36afdc1e10ea4f5c82895b2fe103002"
30
- },
31
- {
32
- "file": "assets/decision-expanded-old_core-600px.png",
33
- "bytes": 167644,
34
- "sha256": "f1586b4bed9423dba1ce138a343ba35de7f95a477aa0946031e99f8c8950279a"
35
- },
36
- {
37
- "file": "assets/decision-expanded-old_core.pdf",
38
- "bytes": 44396,
39
- "sha256": "59aee075c911796b81a32397011acc2de4b9825e161587dd36789f4cb99ceeab"
40
- },
41
- {
42
- "file": "assets/decision-expanded-old_core.png",
43
- "bytes": 312207,
44
- "sha256": "e662de90db73e029550f5badd9a32cc2adb2e3186eb96b12940994bb3ed9eac1"
45
- },
46
- {
47
- "file": "assets/decision-expanded-old_core.svg",
48
- "bytes": 79821,
49
- "sha256": "b63bef1e8d30d9ba15e04778fdff224a70d34606e552619f05c548f221cdfe04"
50
- },
51
- {
52
- "file": "assets/decision-expanded-overview-600px.png",
53
- "bytes": 94909,
54
- "sha256": "f3cab743b7f82cac19857540dd773cdc25ee4db1a5f2b2671e20ff1c1a18ad19"
55
- },
56
- {
57
- "file": "assets/decision-expanded-overview.pdf",
58
- "bytes": 41559,
59
- "sha256": "3734c6df527eed3988a8b14643f225c7c733a3508267305d9681ab3be26da77b"
60
- },
61
- {
62
- "file": "assets/decision-expanded-overview.png",
63
- "bytes": 187333,
64
- "sha256": "c671be3a92da5bb920864d7b0cdbd6ec22e24213d66052e165c20363b319e7e5"
65
- },
66
- {
67
- "file": "assets/decision-expanded-overview.svg",
68
- "bytes": 48619,
69
- "sha256": "87f8f3cf3422548d540b52a943b466497db5eab935d431c36860ad442b059ef4"
70
- },
71
- {
72
- "file": "assets/decision-expanded-ranking-600px.png",
73
- "bytes": 83252,
74
- "sha256": "ce7dcccc65def38efb9822363d6be738e7ff4faaaa7d7a81242f8a42b851edc0"
75
- },
76
- {
77
- "file": "assets/decision-expanded-ranking.pdf",
78
- "bytes": 37133,
79
- "sha256": "e21b19bdb952274ea9e258f5616136195bc72693709592fbf723187bf5e9dba8"
80
- },
81
- {
82
- "file": "assets/decision-expanded-ranking.png",
83
- "bytes": 186197,
84
- "sha256": "e098abf598e3aa51dbef517ed7aeb6c2f7d9312d13bef37d5dbd40ce1dd76aa4"
85
- },
86
- {
87
- "file": "assets/decision-expanded-ranking.svg",
88
- "bytes": 38118,
89
- "sha256": "e327283248812f9104487397ec201b28eef4f1172b058a16286ad224d884ba71"
90
- },
91
- {
92
- "file": "assets/decision-expanded-v3_core-600px.png",
93
- "bytes": 168294,
94
- "sha256": "0630ac755fbfb100ebb3ecb965a5ebe34f511384b1a4c98468d6f0fdbc239d13"
95
- },
96
- {
97
- "file": "assets/decision-expanded-v3_core.pdf",
98
- "bytes": 44466,
99
- "sha256": "426b16796dd3efaf4dc9e50ffd081ec0707831e391629b0342734b3ee334d28e"
100
- },
101
- {
102
- "file": "assets/decision-expanded-v3_core.png",
103
- "bytes": 300404,
104
- "sha256": "d5eaa30f40e03e06414650c023c4994c13182a6847ec423155a9e126df1d16a0"
105
- },
106
- {
107
- "file": "assets/decision-expanded-v3_core.svg",
108
- "bytes": 79901,
109
- "sha256": "5797fa75fd89e31ec23bd47096258e976f916ab5b1adf2f6cdf19f25088f97a1"
110
- },
111
- {
112
- "file": "assets/decision-expanded-v4-600px.png",
113
- "bytes": 76655,
114
- "sha256": "bbd20f118917a754a4844eae746c19e04e07a9629b4ecc8a1dc5e9e83b28c377"
115
- },
116
- {
117
- "file": "assets/decision-expanded-v4.pdf",
118
- "bytes": 37682,
119
- "sha256": "e2015bc6c8b661c58fe52a3f575dab293c77050ab68145032ccaca537ae18ee8"
120
- },
121
- {
122
- "file": "assets/decision-expanded-v4.png",
123
- "bytes": 152599,
124
- "sha256": "47f5d01cea0c8d620f9076f8972bf3457c7e234c99ddcb734250ea3b118e78a9"
125
- },
126
- {
127
- "file": "assets/decision-expanded-v4.svg",
128
- "bytes": 42942,
129
- "sha256": "2ba77ba12ca9eb71105b2e09967a46c223235d2aa82c3175cc0982c7db8cb5b3"
130
- },
131
- {
132
- "file": "assets/decision-expanded-v5-600px.png",
133
- "bytes": 95077,
134
- "sha256": "8eebb01aa741a9cc505e3e3f1841e68d443208efe98b7f9297d7514d1993cdba"
135
- },
136
- {
137
- "file": "assets/decision-expanded-v5.pdf",
138
- "bytes": 40182,
139
- "sha256": "33849a46b149838e4449bc4db80926d6b13b31539884056695ae3235bb5c3190"
140
- },
141
- {
142
- "file": "assets/decision-expanded-v5.png",
143
- "bytes": 187696,
144
- "sha256": "057fe35e93ebe41f3bb8653392c02de27616aa248fb6e2a25026f3d2b1ba9f5a"
145
- },
146
- {
147
- "file": "assets/decision-expanded-v5.svg",
148
- "bytes": 48613,
149
- "sha256": "42f77533052f02089f27f557dee4a6990c835e36da091f4e5458198e496232b6"
150
- },
151
- {
152
- "file": "assets/decision-family-header.png",
153
- "bytes": 2962868,
154
- "sha256": "213511289ce8df038d938ac470e803c427ed57f0f85cc397dd4d79964b866541"
155
- },
156
- {
157
- "file": "assets/decision-question-scaling-600px.png",
158
- "bytes": 40825,
159
- "sha256": "6280e56c1af9f318def3fc13610555cfdc9b9a4af2c69aec8b4226b3cfa41963"
160
- },
161
- {
162
- "file": "assets/decision-question-scaling.pdf",
163
- "bytes": 31467,
164
- "sha256": "7e24268f9abf76a8fd431010f181dc43c5bc09478dbfe75b362435a9ca34cf8c"
165
- },
166
- {
167
- "file": "assets/decision-question-scaling.png",
168
- "bytes": 96655,
169
- "sha256": "b5f48c467a8c9b636a7bbfc34946ffe0f5c8e06210dcb40da706f62e09757910"
170
- },
171
- {
172
- "file": "assets/decision-question-scaling.svg",
173
- "bytes": 25016,
174
- "sha256": "f17f3d7dae22d2dc2ed61cb356e44130dd07b4f2544723e24d291538c8bd5e4f"
175
- },
176
- {
177
- "file": "assets/readout.png",
178
- "bytes": 285481,
179
- "sha256": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085"
180
- },
181
- {
182
- "file": "metrics/comparator-coverage.json",
183
- "bytes": 24628,
184
- "sha256": "fb4e6c8deee00798056aff95ceaaec4f074effbd8b84c20e69667869b1a95c03"
185
- },
186
- {
187
- "file": "metrics/expanded-quality.json",
188
- "bytes": 64295,
189
- "sha256": "7b8936b43bce9d19d9010ab61225a6460160b3aa76c37f5d6df5ba650b433c6c"
190
- },
191
- {
192
- "file": "metrics/figure-evidence.json",
193
- "bytes": 8746,
194
- "sha256": "d9b425316f0d4c1caebb2587cb1027597d421853a089a94e330d016901c5230c"
195
- },
196
- {
197
- "file": "metrics/materials-provenance.json",
198
- "bytes": 1333,
199
- "sha256": "3f052b894ab0780825ef698556483a5900d1e024cfcfaa1e89919261d812adb4"
200
- },
201
- {
202
- "file": "metrics/question-scaling.json",
203
- "bytes": 9034,
204
- "sha256": "a48d1f85555426edecfcfc3a6988c5a87c4f53ef5ec10e86fd489bba50af4191"
205
- },
206
- {
207
- "file": "metrics/semantic-consistency.json",
208
- "bytes": 20784,
209
- "sha256": "5a389236a22e1068be862b2daf88a5caf5a384a9260f0cafbaca9a37b18d8d26"
210
- },
211
- {
212
- "file": "metrics/uncalibrated-quality.json",
213
- "bytes": 19925,
214
- "sha256": "2bc4aca25ef6468bbf3d872e0d72d50476c921717b66ad8b2469f70b051ae735"
215
- },
216
- {
217
- "file": "model-card-example.json",
218
- "bytes": 3775,
219
- "sha256": "9d17a7495a0986d4850a966e282cd3d9e0a31130bdf08c0a9e60461795ce5502"
220
- }
221
- ],
222
- "post_render_finalizer_sha256": "4818899220640baf74e6f87915dfe4ca5f77aedeca40e63c6f95ff9db1cf3d3e",
223
- "linear_latency_finalizer_sha256": "750bfff4297d24237896ef8080758045bc2f12c5fb7ce1a6a790fe871371c161"
224
  }
 
1
  {
2
+ "status": "reviewable-docs-only-payload",
3
+ "family": "Nox",
4
+ "gate_sha256": "687163d21bc76a3e096ade8860b98e55711692e7890c9f5fe263caba1352d3e9",
5
+ "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
6
+ "source_sha256": "92b0b306bfcb8f04c429687b6154695f6b058d168fed30eeae7c96c860bc991f",
7
+ "quality_method": "outcome-informed product weights; unchanged model predictions",
8
+ "all14_models": true,
9
+ "all54_tasks": true,
10
+ "latency_current_API_six_loads_30_samples": true,
11
+ "model_or_runtime_changes": false,
12
+ "public_secret_scan": "PASS",
13
+ "files": {
14
+ "DIAGNOSTICS.md": "5492ed49869e45e45650c2501d02a78f9a2fbc17267c54f90cf034a01ebda2b3",
15
+ "EVALUATION.md": "3866dea6ae78b09bf93ad7b843d067960493217ae84b6ded87bfdac97b424608",
16
+ "QUESTION-SCALING.md": "898f3554434284ff506dbc63c34859750f07b33dcf71f5b6d4ec7a798a56eb28",
17
+ "README.md": "108a5727cda42f93dcc2a3d5fbe653cc42a5fccbe1574ad940553aafe2f1741d",
18
+ "SENSITIVITY.md": "e03ec18140d5c18e606da23cc5b26ef41176ab38019912d85050677e4637f0a0",
19
+ "TASKS.md": "e87c5271049f09ce23fcb7fdcbaa9f03e5c50039b6d7a4005a6b4c7fbc1d30ca",
20
+ "assets/architecture.png": "8fcfa10b3941ee4aec1f5fba20a25e382a7d5de83074071c9ef68e627996de40",
21
+ "assets/architecture.svg": "215d9af6f94cc24847fc1d23a0b2a287b36afdc1e10ea4f5c82895b2fe103002",
22
+ "assets/decision-matrix.pdf": "d1651d91db893ffbdf3261ec476241eb03c18c7a769d19978c668d60a1ac844f",
23
+ "assets/decision-matrix.png": "b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc",
24
+ "assets/decision-matrix.svg": "dc92d1e3f3cf0518ee44f1bc512aae74732e1db82b919774a52c0da7c744b0b4",
25
+ "assets/decision-question-scaling.pdf": "d756f60c142ecaf647e59ea5bb3a611a39d1572a6471cd49afb7ef602a23fb3b",
26
+ "assets/decision-question-scaling.png": "722e664974b28324693f5f4af6609330c278385ac118ed6f757b8fc0afec64ac",
27
+ "assets/decision-question-scaling.svg": "cc591eafa335b5d16edd306d6368679afd283a4b14d5f83fc77780a352b1d315",
28
+ "assets/decision-ranking.pdf": "f3e037ada6c30871d9ebe0087fa3d122346f7ab1c927515878d524c492b7799a",
29
+ "assets/decision-ranking.png": "1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0",
30
+ "assets/decision-ranking.svg": "c8de982ba51fb9d2b729191018220964426a0f524a00a796553b8879db271d1a",
31
+ "assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
32
+ "metrics/benchmark.json": "50dac7884849cbdb8be58c1c82b811f51543a08a548253cb8295f4e6444c4718",
33
+ "metrics/evaluation-provenance.json": "1f9c8b1e03d521160a8426798394418748daa72014dfc7272578d1e83fe237b4",
34
+ "metrics/question-scaling.json": "9407f7ebf86581cdea21e704cdbebb9ced0a485eb980925f48233ebea0f830dc"
35
+ }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
  }
QUESTION-SCALING.md CHANGED
@@ -1,42 +1,20 @@
1
  # Request latency
2
 
3
- Current published Nox v1.3, measured on one otherwise idle AMD gfx942 GPU.
4
-
5
- Three independently loaded blocks provide 30 measured requests per point after three warmups per block. End-to-end Python latency includes tokenization and inference, excluding model loading and network. No cross-request prefix cache is enabled. These are sequential requests, not concurrent-service throughput.
6
 
7
  ![Question scaling](assets/decision-question-scaling.png)
8
 
9
- ## Choice
10
-
11
- | Questions | p50 ms ↓ | p95 ms ↓ | Input tokens |
12
- |---:|---:|---:|---:|
13
- | 1 | 32.732 | 33.474 | 309 |
14
- | 2 | 33.319 | 33.560 | 618 |
15
- | 4 | 41.779 | 42.215 | 1236 |
16
- | 8 | 69.642 | 70.728 | 2472 |
17
- | 16 | 141.020 | 141.853 | 4944 |
18
- | 32 | 282.028 | 283.684 | 9888 |
19
-
20
- ## Noul
21
-
22
- | Questions | p50 ms ↓ | p95 ms ↓ | Input tokens |
23
  |---:|---:|---:|---:|
24
- | 1 | 32.685 | 33.001 | 262 |
25
- | 2 | 32.823 | 33.581 | 524 |
26
- | 4 | 38.836 | 39.629 | 1048 |
27
- | 8 | 64.279 | 65.077 | 2096 |
28
- | 16 | 129.265 | 130.600 | 4192 |
29
- | 32 | 259.156 | 261.597 | 8384 |
30
 
31
- ## Score
32
 
33
- | Questions | p50 ms ↓ | p95 ms ↓ | Input tokens |
34
- |---:|---:|---:|---:|
35
- | 1 | 33.319 | 34.269 | 307 |
36
- | 2 | 33.356 | 33.767 | 614 |
37
- | 4 | 41.859 | 42.210 | 1228 |
38
- | 8 | 69.734 | 70.124 | 2456 |
39
- | 16 | 141.065 | 141.961 | 4912 |
40
- | 32 | 282.330 | 285.689 | 9824 |
41
 
42
- The bundled runtime, normalization profile and temperature belong to this exact measured release. The source experiment used a counterbalanced matched design; this display contains only the current release. No cross-model latency conclusion is inferred.
 
1
  # Request latency
2
 
3
+ Nox · distinct Choice questions with **499 input tokens per question**. Only the number of questions changes.
 
 
4
 
5
  ![Question scaling](assets/decision-question-scaling.png)
6
 
7
+ | Questions | p50 ms ↓ | p95 ms ↓ | Peak allocated GiB |
 
 
 
 
 
 
 
 
 
 
 
 
 
8
  |---:|---:|---:|---:|
9
+ | 1 | 32.678 | 33.719 | 8.079 |
10
+ | 2 | 36.370 | 36.647 | 8.146 |
11
+ | 4 | 60.415 | 60.952 | 8.280 |
12
+ | 8 | 103.619 | 104.181 | 8.549 |
13
+ | 16 | 206.991 | 207.624 | 8.549 |
14
+ | 32 | 414.199 | 415.371 | 8.549 |
15
 
16
+ Six independently loaded process blocks supply 30 measured requests per point after warmup. End-to-end Python request latency includes rendering, tokenization, inference, output assembly and final synchronization. Model loading and network are excluded. Requests are sequential, with no cross-request prefix cache.
17
 
18
+ The measured implementation is bound byte-for-byte to the current published API, verified by an offline Hub-download proof and full 3,160-answer regression. This documentation update changes no inference code or weights. [Exact measurements and immutable runtime identity](metrics/question-scaling.json).
 
 
 
 
 
 
 
19
 
20
+ These fixed short-input Choice measurements do not establish concurrent HTTP throughput, long-context scaling, other question-type performance or a cross-hardware speed ranking.
README.md CHANGED
@@ -14,13 +14,11 @@ tags:
14
  - rocm
15
  ---
16
 
17
- ![Decision 1.0 — Your move.](assets/decision-family-header.png)
18
-
19
  # Decision-1.0-Nox
20
 
21
  *Nox, Latin for night.*
22
 
23
- **Your move.** Give Nox a state, questions and possible answers. It returns decisions and probabilities with labels defined at runtime.
24
 
25
  **4.208B parameters · 16K complete-question budget · English / Chinese evaluated · Apache 2.0**
26
 
@@ -34,59 +32,42 @@ tags:
34
 
35
  ## Measured capability
36
 
37
- **75.03% four-panel mean** across 2,720 decisions. Results include strengths and remaining gaps across all 27 tasks.
38
-
39
- | Model | Mean accuracy ↑ | Decisions | Composition | Reading | Inference |
40
- |---|---:|---:|---:|---:|---:|
41
- | Nox · 4B · v1.3 | **75.03** | **83.00** | **51.79** | 79.06 | **86.25** |
42
- | Sol · 2B · v1.3 | 70.14 | 73.75 | 46.08 | 76.56 | 84.17 |
43
- | Kev · 9B | 73.09 | 76.36 | 45.54 | 86.72 | 83.75 |
44
- | Decider · 2B | 71.75 | 64.01 | 46.58 | 92.03 | 84.38 |
45
- | Qwen3.5 · 4B · untuned | 70.25 | 69.89 | 43.33 | 87.97 | 79.79 |
46
- | Qwen3.5 · 2B · untuned | 60.54 | 57.12 | 39.00 | 73.75 | 72.29 |
47
- | Laya · Upstream default | 52.44 | 57.01 | 37.75 | 51.25 | 63.75 |
48
- | Jev · 1.13.0 · frontier | 82.45 | 79.10 | 66.38 | 94.53 | 89.79 |
49
-
50
- Accuracy (%). Each panel contributes one quarter with its frozen family/source weights. **Bold** marks Sol or Nox strictly above every external open reference for that metric; the other Decision model and Jev are excluded from the threshold. Jev is a closed-service frontier reference. [Methods and uncertainty](EVALUATION.md).
51
-
52
- ![Decision comparison](assets/decision-expanded-ranking.png)
53
-
54
- ![Four-panel capability overview](assets/decision-expanded-overview.png)
55
-
56
- ## Detailed capabilities
57
-
58
- ![General decisions](assets/decision-expanded-old_core.png)
59
-
60
- ![Compositional tasks](assets/decision-expanded-v3_core.png)
61
 
62
- ![Natural reading](assets/decision-expanded-v4.png)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
63
 
64
- ![Reading and inference](assets/decision-expanded-v5.png)
65
 
66
- ## Semantic consistency
67
 
68
- | Model | Consistent / eligible groups | Consistency ↑ |
69
- |---|---:|---:|
70
- | Nox · 4B · v1.3 | 174/192 | **90.62** |
71
- | Sol · 2B · v1.3 | 167/192 | **86.98** |
72
- | Kev · 9B | 143/192 | 74.48 |
73
- | Decider · 2B | 126/192 | 65.62 |
74
- | Qwen3.5 · 4B · untuned | 98/192 | 51.04 |
75
- | Qwen3.5 · 2B · untuned | 94/192 | 48.96 |
76
- | Laya · Upstream default | 115/192 | 59.90 |
77
- | Jev · 1.13.0 · frontier | 158/192 | 82.29 |
78
 
79
- Mixed equivalent variants of language, keys and presentation; **consistency is not accuracy**. A consistently wrong answer still counts. This observed-core supplement weights 192 eligible semantic groups equally and is outside the four-panel mean. [Definitions and exact counts](metrics/semantic-consistency.json).
80
 
81
- ## More questions, measured
82
 
83
- ![Current release request latency](assets/decision-question-scaling.png)
84
 
85
- Same inputs and physical AMD gfx942 GPU; 30 measured requests per point across three blocks. Python request latency includes tokenization and inference, excluding loading and network. [p95 and all three native types](QUESTION-SCALING.md).
86
 
87
  ## Try it
88
 
89
- Download `hf download llm-semantic-router/Decision-1.0-Nox --revision v1.3 --local-dir decision-model`, then follow [ROCm setup](RUNTIME.md). In that container, with the model mounted at `/model`:
90
 
91
  ```python
92
  from decision import DecisionModel
@@ -96,7 +77,7 @@ model = DecisionModel.from_pretrained("/model", local_files_only=True)
96
  print(model.decide(**REQUEST)["answers"])
97
  ```
98
 
99
- [Tested request and output](model-card-example.json) · [Install and API guide](USAGE.md)
100
 
101
  The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.
102
 
 
14
  - rocm
15
  ---
16
 
 
 
17
  # Decision-1.0-Nox
18
 
19
  *Nox, Latin for night.*
20
 
21
+ **Your move.** Give Nox a state, questions and possible answers. It returns typed decisions and probabilities, with labels defined at runtime.
22
 
23
  **4.208B parameters · 16K complete-question budget · English / Chinese evaluated · Apache 2.0**
24
 
 
32
 
33
  ## Measured capability
34
 
35
+ **72.84% overall accuracy** across 3,766 scored decisions and 54 tasks. Nox leads Kev-4B by **2.75 percentage points** on this decision-focused comparison; reading and transfer remain opportunities to improve.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
 
37
+ | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
38
+ |---|---:|---:|---:|---:|---:|---:|---:|
39
+ | Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
40
+ | Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
41
+ | Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** |
42
+ | Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
43
+ | Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
44
+ | Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
45
+ | Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 |
46
+ | Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 |
47
+ | Sol | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 |
48
+ | Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 |
49
+ | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
50
+ | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
51
+ | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
52
+ | Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
53
 
54
+ Accuracy (%). Overall weights: Decisions **30%**, Composition **25%**, Reading **15%**, Inference **15%**, Transfer **15%**. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.
55
 
56
+ ![Decision model ranking](assets/decision-ranking.png)
57
 
58
+ ![Capability matrix](assets/decision-matrix.png)
 
 
 
 
 
 
 
 
 
59
 
60
+ [All 54 tasks](TASKS.md) · [Order, missing-evidence and calibration diagnostics](DIAGNOSTICS.md) · [Methods and uncertainty](EVALUATION.md)
61
 
62
+ ## More questions, one request
63
 
64
+ ![Question-count latency](assets/decision-question-scaling.png)
65
 
66
+ Distinct Choice questions at a fixed **499 input tokens per question**. Thirty measurements per point across six independently loaded processes on an otherwise idle AMD gfx942 GPU. Python latency includes tokenization and inference; loading and network are excluded. [p50, p95 and memory](QUESTION-SCALING.md).
67
 
68
  ## Try it
69
 
70
+ Download the current model with `hf download llm-semantic-router/Decision-1.0-Nox --local-dir decision-model`, then follow [ROCm setup](RUNTIME.md). In that container, with the model mounted at `/model`:
71
 
72
  ```python
73
  from decision import DecisionModel
 
77
  print(model.decide(**REQUEST)["answers"])
78
  ```
79
 
80
+ [Tested SystemOne request and output](model-card-example.json) · [Install and typed API guide](USAGE.md)
81
 
82
  The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.
83
 
SENSITIVITY.md ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Product weights were changed after earlier results were observed. These are identical model predictions under different weights, not trained-model improvements.
2
+
3
+ | Model | Current 30/25/15/15/15 | Earlier 25/25/15/15/20 | Original four-panel mean |
4
+ |---|---:|---:|---:|
5
+ | Jev | 81.05 | 81.45 | 82.45 |
6
+ | Lux | 76.72 | 76.43 | 79.06 |
7
+ | Nox | 72.84 | 72.09 | 75.03 |
8
+ | Kev-9B | 71.89 | 72.01 | 73.19 |
9
+ | Kev-4B | 70.09 | 70.30 | 71.73 |
10
+ | Qwen3.5-9B | 69.73 | 69.70 | 71.99 |
11
+ | Decider | 67.71 | 67.97 | 71.75 |
12
+ | Qwen3.5-4B | 67.29 | 67.24 | 70.25 |
13
+ | Sol | 66.32 | 65.48 | 70.14 |
14
+ | Kev-0.8B | 58.28 | 58.33 | 59.75 |
15
+ | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
16
+ | Laya · English | 51.03 | 50.85 | 51.76 |
17
+ | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
18
+ | Kai | 46.49 | 46.80 | 48.16 |
TASKS.md ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # All 54 tasks
2
+
3
+ Accuracy (%) on the same requested rows. Bold marks a Decision-family result strictly above every external open or untuned reference; Jev and the other Decision models do not set that threshold. Ties are not bold. Each task uses its full requested denominator. These task rows are diagnostic; the headline retains its declared within-panel weights.
4
+
5
+ <details>
6
+ <summary>Decisions · 10 tasks</summary>
7
+
8
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
9
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
10
+ | News classification | 128 | 85.16 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 | 78.12 |
11
+ | Boolean constraints | 64 | 100.00 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 | 53.12 |
12
+ | Entity classification | 112 | 96.43 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 | 50.00 |
13
+ | Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 | 81.25 |
14
+ | Evidence placement | 96 | 30.21 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 | 1.04 |
15
+ | Ordered rubric | 64 | 100.00 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 | 6.25 |
16
+ | Relation composition | 96 | 56.25 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 | 51.04 |
17
+ | Scoped evidence | 96 | 89.58 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 | 29.17 |
18
+ | State tracking | 96 | 33.33 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 | 30.21 |
19
+ | In / out of menu | 64 | 100.00 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 | 42.19 |
20
+
21
+ </details>
22
+
23
+ <details>
24
+ <summary>Composition · 10 tasks</summary>
25
+
26
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
27
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
28
+ | Record identity | 80 | 76.25 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 | 51.25 |
29
+ | Capacity assignment | 80 | 68.75 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 | 50.00 |
30
+ | Constraint assignment | 80 | 55.00 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 | 31.25 |
31
+ | Intent routing · EN | 120 | 91.67 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 | 72.50 |
32
+ | Intent routing · ZH | 120 | 88.33 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 | 67.50 |
33
+ | Multiset reconciliation | 80 | 58.75 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 | 32.50 |
34
+ | Ordered service loss | 80 | 42.50 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 | 21.25 |
35
+ | Conflicting rule closure | 80 | 66.25 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 | 23.75 |
36
+ | Temporal exclusion | 80 | 41.25 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 | 22.50 |
37
+ | Transaction recovery | 80 | 75.00 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 | 26.25 |
38
+
39
+ </details>
40
+
41
+ <details>
42
+ <summary>Reading · 3 tasks</summary>
43
+
44
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
45
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
46
+ | Yes / no reading | 160 | 92.50 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 | 74.38 |
47
+ | Reading · EN | 160 | 96.88 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 | 39.38 |
48
+ | Reading · ZH | 160 | 96.25 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 | 35.62 |
49
+
50
+ </details>
51
+
52
+ <details>
53
+ <summary>Inference · 4 tasks</summary>
54
+
55
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
56
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
57
+ | Contextual reasoning | 120 | 86.67 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 | 32.50 |
58
+ | Answerability | 120 | 90.83 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 | 55.83 |
59
+ | Textual entailment | 120 | 82.50 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 | 34.17 |
60
+ | Scientific inference | 120 | 99.17 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 | 95.83 |
61
+
62
+ </details>
63
+
64
+ <details>
65
+ <summary>Transfer · 27 tasks</summary>
66
+
67
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
68
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
69
+ | Buried emotion | 20 | 70.00 | 40.00 | 20.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 | 55.00 |
70
+ | Buried paraphrase | 20 | 95.00 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 | 35.00 |
71
+ | Buried entailment | 20 | 95.00 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 | 55.00 |
72
+ | Buried offensive-language detection | 20 | 55.00 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 | 75.00 |
73
+ | Combined policy conditions | 32 | 96.88 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 | 34.38 |
74
+ | Policy exceptions | 32 | 100.00 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 | 62.50 |
75
+ | Policy negation | 32 | 90.62 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 | 56.25 |
76
+ | Authorization contrast | 40 | 100.00 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 | 50.00 |
77
+ | Deadline contrast | 40 | 92.50 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 | 47.50 |
78
+ | Emotion | 80 | 67.50 | 61.25 | 40.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 | 52.50 |
79
+ | MMLU | 80 | 88.75 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 | 31.25 |
80
+ | MMLU-Pro | 200 | 84.00 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 | 17.50 |
81
+ | Paraphrase | 80 | 87.50 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 | 47.50 |
82
+ | Question entailment | 80 | 91.25 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 | 71.25 |
83
+ | Science questions | 80 | 100.00 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 | 92.50 |
84
+ | Offensive-language detection | 80 | 76.25 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 | 78.75 |
85
+ | Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 | 50.00 |
86
+ | Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 | 50.00 |
87
+ | Evidence control · deadline | 10 | 100.00 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 | 50.00 |
88
+ | Evidence control · late fee | 10 | 100.00 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 | 30.00 |
89
+ | Evidence control · quantity limit | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 | 50.00 |
90
+ | Evidence control · return window | 10 | 50.00 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 | 50.00 |
91
+ | Evidence control · shipping delay | 10 | 90.00 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 | 20.00 |
92
+ | Evidence control · sla response | 10 | 100.00 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 | 20.00 |
93
+ | Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 | 50.00 |
94
+ | Evidence control · volume discount | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 | 30.00 |
95
+ | Evidence control · warranty claim | 10 | 90.00 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 | 10.00 |
96
+
97
+ </details>
assets/decision-matrix.pdf ADDED
Binary file (28.6 kB). View file
 
assets/decision-matrix.png ADDED

Git LFS Details

  • SHA256: b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc
  • Pointer size: 131 Bytes
  • Size of remote file: 345 kB
assets/decision-matrix.svg ADDED
assets/decision-question-scaling.pdf CHANGED
Binary files a/assets/decision-question-scaling.pdf and b/assets/decision-question-scaling.pdf differ
 
assets/decision-question-scaling.png CHANGED

Git LFS Details

  • SHA256: b5f48c467a8c9b636a7bbfc34946ffe0f5c8e06210dcb40da706f62e09757910
  • Pointer size: 130 Bytes
  • Size of remote file: 96.7 kB

Git LFS Details

  • SHA256: 722e664974b28324693f5f4af6609330c278385ac118ed6f757b8fc0afec64ac
  • Pointer size: 131 Bytes
  • Size of remote file: 104 kB
assets/decision-question-scaling.svg CHANGED
assets/decision-ranking.pdf ADDED
Binary file (24.6 kB). View file
 
assets/decision-ranking.png ADDED

Git LFS Details

  • SHA256: 1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0
  • Pointer size: 131 Bytes
  • Size of remote file: 226 kB
assets/decision-ranking.svg ADDED
metrics/benchmark.json ADDED
The diff for this file is too large to render. See raw diff
 
metrics/evaluation-provenance.json ADDED
@@ -0,0 +1,1019 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "protocol": {
3
+ "version": "decision-priority-benchmark-v4",
4
+ "utc": "2026-09-22T05:05:12.118255+00:00",
5
+ "status": "User-requested product weighting, observed regression benchmark; release remains evidence-gated",
6
+ "authorization": "Latest user explicitly requested transfer20%\u219215% and assigned released5% to the stronger original-decision/composition area. Original decisions chosen from already observed same-size gaps. This is outcome-informed product weighting, not neutral/blind prospective benchmark design.",
7
+ "weights": {
8
+ "old_core": "3/10",
9
+ "v3_core": "1/4",
10
+ "v4": "3/20",
11
+ "v5": "3/20",
12
+ "transfer_v9_test": "3/20"
13
+ },
14
+ "accuracy_mean": "Exact weighted categorical accuracy; originalwithin-panelweights and requested denominators unchanged.",
15
+ "original_panel_registry_sha256": "1aefdf58d56e832e765194d0fa53a8647acea7d2ad2fc25ac95380bf85949fd2",
16
+ "transfer_protocol_sha256": "f0c0f8982bfbb9edfc5e486f682906612fae579f28c58c2fa656ad9d10874d16",
17
+ "transfer_payload_sha256": "97d82909bcbd8756904d6ed59b31e16d16086d62990ccdc981dc95b8c274e0c5",
18
+ "transfer_labels_sha256": "0b64882d58e2ac1e0f5e2e102f916907554c94c07f2774b8c092683cfc81973c",
19
+ "transfer_accuracy_questions": 1046,
20
+ "original_core_questions": 2720,
21
+ "diagnostics": "Preserve internal-v2 source-generalization/order/unknown/proper-score/efficiency axes separately; no millisecond or ECE averaging into accuracy.",
22
+ "roster": {
23
+ "Nox": {
24
+ "role": "Decision family",
25
+ "size": "4B",
26
+ "repo": "llm-semantic-router/Decision-1.0-Nox"
27
+ },
28
+ "Sol": {
29
+ "role": "Decision family",
30
+ "size": "2B",
31
+ "repo": "llm-semantic-router/Decision-1.0-Sol"
32
+ },
33
+ "Kai": {
34
+ "role": "Decision family",
35
+ "size": "0.572B encoder",
36
+ "repo": "llm-semantic-router/Decision-1.0-Kai",
37
+ "revision": "2079354070d02f4c2b1e60bce6b99f60dfad9622",
38
+ "contract": "shipped1024complete-token admission; FP32; defaultB8; no limit override"
39
+ },
40
+ "Lux": {
41
+ "role": "Decision family",
42
+ "size": "9B",
43
+ "repo": "llm-semantic-router/Decision-1.0-Lux",
44
+ "pending_unreleased": true
45
+ },
46
+ "kev-0.8b": {
47
+ "role": "open reference",
48
+ "size": "0.8B",
49
+ "repo": "jaredpalmer/kev-0.8b",
50
+ "revision": "54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8"
51
+ },
52
+ "kev-4b": {
53
+ "role": "open reference",
54
+ "size": "4B",
55
+ "repo": "jaredpalmer/kev-4b",
56
+ "revision": "485ace8703592fcf405488b262449990824cfed1"
57
+ },
58
+ "kev-9b": {
59
+ "role": "open reference",
60
+ "size": "9B",
61
+ "repo": "jaredpalmer/kev-9b",
62
+ "revision": "2629c06a5aeb0feb3b9783bafed17ed8f39ecf5c"
63
+ },
64
+ "Decider": {
65
+ "role": "open reference",
66
+ "size": "2B",
67
+ "repo": "Mapika/decider-2b"
68
+ },
69
+ "Laya-base": {
70
+ "role": "open reference",
71
+ "size": "0.421B encoder",
72
+ "repo": "convaiinnovations/laya",
73
+ "mode": "english"
74
+ },
75
+ "Laya-multilingual": {
76
+ "role": "open reference",
77
+ "size": "0.322B encoder",
78
+ "repo": "convaiinnovations/laya",
79
+ "mode": "multilingual"
80
+ },
81
+ "Jev": {
82
+ "role": "closed frontier",
83
+ "size": "undisclosed",
84
+ "served_identity": "jev-1.13.0"
85
+ },
86
+ "Qwen3.5-2B": {
87
+ "role": "untuned reference",
88
+ "size": "2B",
89
+ "repo": "Qwen/Qwen3.5-2B",
90
+ "adapter": "original locked chat/LM-head letter readout; no training"
91
+ },
92
+ "Qwen3.5-4B": {
93
+ "role": "untuned reference",
94
+ "size": "4B",
95
+ "repo": "Qwen/Qwen3.5-4B",
96
+ "adapter": "original locked chat/LM-head letter readout; no training"
97
+ },
98
+ "Qwen3.5-9B": {
99
+ "role": "untuned reference",
100
+ "size": "9B",
101
+ "repo": "Qwen/Qwen3.5-9B",
102
+ "adapter": "original locked chat/LM-head letter readout; no training"
103
+ }
104
+ },
105
+ "internal_extra_references": [
106
+ "LLM2Jev2B",
107
+ "LLM2Jev4B",
108
+ "Nimble9B"
109
+ ],
110
+ "cohort_gates": {
111
+ "Sol": {
112
+ "required_same_size_open": [
113
+ "Decider"
114
+ ],
115
+ "additional_measured_internal_peer": [
116
+ "LLM2Jev2B"
117
+ ],
118
+ "required_same_size_untuned": [
119
+ "Qwen3.5-2B"
120
+ ]
121
+ },
122
+ "Nox": {
123
+ "required_same_size_open": [
124
+ "kev-4b"
125
+ ],
126
+ "additional_measured_internal_peer": [
127
+ "LLM2Jev4B"
128
+ ],
129
+ "required_same_size_untuned": [
130
+ "Qwen3.5-4B"
131
+ ]
132
+ },
133
+ "Lux": {
134
+ "must_exceed_current_family": [
135
+ "Nox",
136
+ "Sol"
137
+ ],
138
+ "required_same_size_open": [
139
+ "kev-9b"
140
+ ],
141
+ "additional_measured_internal_peer": [
142
+ "Nimble9B"
143
+ ],
144
+ "required_same_size_untuned": [
145
+ "Qwen3.5-9B"
146
+ ]
147
+ }
148
+ },
149
+ "publication": {
150
+ "first_expanded_cards": "Nox and Sol independently eligible only when their own required same-size comparisons are complete and strictly beaten; table includes complete available public roster. Unreleased Lux need not delay either card.",
151
+ "later_updates": "Require newaggregate strictly higher than the current published same model plus maintained same-size lead; show regressions and full diagnostic metrics honestly.",
152
+ "Lux_first_release": "Newweightedaggregate strictly exceeds then-currentNox/Sol and known same-sizeopen references; old4panelgate superseded before heldout/publication.",
153
+ "missing_scores": "Never fill from different benchmark, revision, successful-only denominator or inferred model size.",
154
+ "accuracy_uncertainty": "Report paired component bootstrap intervals; no new positive-CI gate invented.",
155
+ "versions": "No trainingversion labels in visiblecard/table/figures; retain immutable revision/weight/adapter identities in machine-readable provenance.",
156
+ "rank": "All public roster descending actual score. Decision family copper; distinct restrained colors for Kev, Decider, Laya, frontierJev and untunedQwen. White background, no logos or trainingversion labels.",
157
+ "matrix": "Minimal readable typography, full task coverage, no logo; labels omit trainingversions.",
158
+ "original_comparison": "Retain original four-panel and prior five-panel25/25/15/15/20 comparisons in methods/provenance. Explicit outcome-informed product-priority amendment; no universal or prospective performance claim."
159
+ },
160
+ "training_integrity": "Frozen ongoing training/SELECT/CAL unchanged. Testresults never select anothercheckpoint; future targetedtraining uses independent TRAIN/SELECT/CAL. Observed tests labeledregression.",
161
+ "supersedes": {
162
+ "protocol": {
163
+ "path": "eval/decision-benchmark-v3/PROTOCOL.json",
164
+ "sha256": "d0b7c667fd4de0dfa5268d2d83aebf40e2b0c43d615ad77f4c2264e35a137534"
165
+ },
166
+ "prior_full_statistics": {
167
+ "path": "analysis/decision-benchmark-v3-results/full01/scored/STATISTICS.json",
168
+ "sha256": "aaed8ca33a874300545b81566cb0ba37bf7cc8fcaac954beb9094aeef8ad82f7"
169
+ }
170
+ },
171
+ "sensitivity": "Retain full25/25/15/15/20 results alongside30/25/15/15/15 in methods/provenance. Reweighting gains are never described as training improvements. No model selection, calibration, predictions or denominators changed."
172
+ },
173
+ "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
174
+ "manifest_sha256": "3199a40b2719cad9b1e815629cf5ebad923a9a399188c1199a9d27e918bda9e3",
175
+ "qualified_runtime": {
176
+ "python": "3.12.13",
177
+ "numpy": "2.3.5"
178
+ },
179
+ "source_sha256": "b00c60f21ace9bd258a1a79673b8cdbce3101ada2fed5547811f70bf030fcd91",
180
+ "model_manifest": {
181
+ "Nox": {
182
+ "panels": {
183
+ "old_core": {
184
+ "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/old_core/normalized.jsonl",
185
+ "sha256": "9b6b7db907e8e458cd33982c0f7c45e5052c3248285d927990f160d5570c8976"
186
+ },
187
+ "v3_core": {
188
+ "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v3_core/normalized.jsonl",
189
+ "sha256": "49575aef2261c191e289d781857664c57fd6181701c03f02515292f2fa31b40c"
190
+ },
191
+ "v4": {
192
+ "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v4/normalized.jsonl",
193
+ "sha256": "9d85618b1a05c2ed68f6a188dadf19a64d6ede6d5a1b0ef58afa27a13906238b"
194
+ },
195
+ "v5": {
196
+ "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v5/normalized.jsonl",
197
+ "sha256": "5eb7faaa45a22dad414af1363d6087b8c379c86599df3c14483a3a648ca6bfd5"
198
+ }
199
+ },
200
+ "transfer_rows": {
201
+ "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Nox/v9-transfer-v9-test-published-temperature-rows.json",
202
+ "sha256": "77b92a18433b260ac0e6f4e0b6c9ec3a94e16aaed870a4078b1d7f32ef18b169"
203
+ },
204
+ "evidence": [
205
+ {
206
+ "path": "results/expanded/nox-retention-v1/QUALITY-COLLECTION.json",
207
+ "sha256": "7608258282ec1e9731e4d2ea4dee7206fda8900055dc64f9193d2566172ca86a"
208
+ },
209
+ {
210
+ "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Nox/REPORT.json",
211
+ "sha256": "8e78910746c9e78862b3aadd94490054a681f3ad013b43ec19bfbc2434111ed7"
212
+ }
213
+ ]
214
+ },
215
+ "Sol": {
216
+ "panels": {
217
+ "old_core": {
218
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/old_core/normalized.jsonl",
219
+ "sha256": "5fc0b3e49a5de576f7621839130855c178b12d942dc357aaf6b4dd9735e0a6a8"
220
+ },
221
+ "v3_core": {
222
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v3_core/normalized.jsonl",
223
+ "sha256": "b2c99fbd739b849a66a55929a658de4517a7ffaa4496255566e6972ba28d4f15"
224
+ },
225
+ "v4": {
226
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v4/normalized.jsonl",
227
+ "sha256": "3c28302750e02d4cead0aacd66a323734e773f726c4c41db658dd6461251b28b"
228
+ },
229
+ "v5": {
230
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v5/normalized.jsonl",
231
+ "sha256": "6ddb0496012b596ea122afa4db4c4158e57dfb934e9ae29eb006756b652e31e1"
232
+ }
233
+ },
234
+ "transfer_rows": {
235
+ "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/v9-transfer-v9-test-published-temperature-rows.json",
236
+ "sha256": "b05799417cb3715de633135c58f713960b75bdb9ed30eb8912a197d33ee3d85b"
237
+ },
238
+ "evidence": [
239
+ {
240
+ "path": "results/expanded/sol-composition-v2/QUALITY-COLLECTION.json",
241
+ "sha256": "9b77aac87b85cbf3a74e86af4a2b9a03fd9e1ed854356a92ff54ccd0d47a16f4"
242
+ },
243
+ {
244
+ "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/REPORT.json",
245
+ "sha256": "5d8ebb339293687cec9089f54fb96aabf3330f16111487a49489f5e52bbb6f29"
246
+ }
247
+ ]
248
+ },
249
+ "kev-0.8b": {
250
+ "panels": {
251
+ "old_core": {
252
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-old_core-normalized.jsonl",
253
+ "sha256": "b492b4512ac164f83a319f31bbc9e086cae12f87ffe2c14f307481ffb240807c"
254
+ },
255
+ "v3_core": {
256
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v3_core-normalized.jsonl",
257
+ "sha256": "70cbfc69dfc72ac0c5ae01305ec63a9f31129c115b7068c6824cd22ead013d68"
258
+ },
259
+ "v4": {
260
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v4-normalized.jsonl",
261
+ "sha256": "761dfd69acbface101f2c21fd27d2db684c5c8b8cc425890a075729ae1f7db24"
262
+ },
263
+ "v5": {
264
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v5-normalized.jsonl",
265
+ "sha256": "51b6fc5f826ca8db44df43885f34d857ddd5b5a66453bbdc1a83fcf488928bdb"
266
+ }
267
+ },
268
+ "transfer_rows": {
269
+ "path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/v9-transfer-v9-test-published-temperature-rows.json",
270
+ "sha256": "d6abed88001aeac630219466f02c1aad89fcbd354b3ae533507f3c2f7fe42898"
271
+ },
272
+ "evidence": [
273
+ {
274
+ "path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/REPORT.json",
275
+ "sha256": "6760dcbaa7d85e6c2e893e4c80ab8e66e0b84d7cfa26ae9e25909f48a05f8de4"
276
+ },
277
+ {
278
+ "path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
279
+ "sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
280
+ }
281
+ ]
282
+ },
283
+ "kev-4b": {
284
+ "panels": {
285
+ "old_core": {
286
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-old_core-normalized.jsonl",
287
+ "sha256": "41dfbb7c9c4381effc219f9123b51bd4d688b3e746cec974ef930730b9dfd6b1"
288
+ },
289
+ "v3_core": {
290
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v3_core-normalized.jsonl",
291
+ "sha256": "2ec54da92c4e72412e89b3c07013b54729f5899ba8542eeadf962b04a1f3a4b8"
292
+ },
293
+ "v4": {
294
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v4-normalized.jsonl",
295
+ "sha256": "88c576c3d62ec11686dd5ebf360971a9087779e45a346897e2badf874a754a12"
296
+ },
297
+ "v5": {
298
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v5-normalized.jsonl",
299
+ "sha256": "d33bbb60134314dcb550060176ed0a3ae91ec0c6b033cceb40e3255b33666898"
300
+ }
301
+ },
302
+ "transfer_rows": {
303
+ "path": "analysis/kev-reciprocal-v1/reports/kev-4b/v9-transfer-v9-test-published-temperature-rows.json",
304
+ "sha256": "c2cb16a936be4c10ca6b1870c69426ddd6b51629c401f71ae1da652b3bd2ac7e"
305
+ },
306
+ "evidence": [
307
+ {
308
+ "path": "analysis/kev-reciprocal-v1/reports/kev-4b/REPORT.json",
309
+ "sha256": "366fb0d27a4d732e5e270c264521d9a93ccabc4f0ddf5b7d3e9f1ecc5f460406"
310
+ },
311
+ {
312
+ "path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
313
+ "sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
314
+ }
315
+ ]
316
+ },
317
+ "kev-9b": {
318
+ "panels": {
319
+ "old_core": {
320
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-old_core-normalized.jsonl",
321
+ "sha256": "a3b11fee4bc5e207cc42e7f4d8dde05df6cfc73d9446ac33cd6015e1ab9e587d"
322
+ },
323
+ "v3_core": {
324
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-v3_core-normalized.jsonl",
325
+ "sha256": "bb90e5701a151272f97caa4107888ba2379969b4e3f22abb4670a528b50d7729"
326
+ },
327
+ "v4": {
328
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-v4-normalized.jsonl",
329
+ "sha256": "b9c50612c0be1b7b5647779d8489c185ea7f02673da176b1d965673da604d644"
330
+ },
331
+ "v5": {
332
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-v5-normalized.jsonl",
333
+ "sha256": "2893714030357718e662633fddf31a328e56a02fc3ae71ed34612f480e50d4ee"
334
+ }
335
+ },
336
+ "transfer_rows": {
337
+ "path": "analysis/kev-reciprocal-v1/reports/kev-9b/v9-transfer-v9-test-published-temperature-rows.json",
338
+ "sha256": "e712bf23acea6a9431001f5424639d1e3738789907ff2e367db7c9dad4e0d239"
339
+ },
340
+ "evidence": [
341
+ {
342
+ "path": "analysis/kev-reciprocal-v1/reports/kev-9b/REPORT.json",
343
+ "sha256": "96dcf11961f510d05c011aa3fd8dfa6043348471b8e800d6bac2d691d0815d54"
344
+ },
345
+ {
346
+ "path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
347
+ "sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
348
+ }
349
+ ]
350
+ },
351
+ "Jev": {
352
+ "panels": {
353
+ "old_core": {
354
+ "path": "eval/heldout/official-predictions.jsonl",
355
+ "sha256": "069e48e1fdf1b5f3d95cc0097e4179ac3faab97bb357b0ce897cfb463b87cbaf",
356
+ "bytes": 185263
357
+ },
358
+ "v3_core": {
359
+ "path": "eval/heldout/v3/official/normalized/core.jsonl",
360
+ "sha256": "54707354bfe5d96b5de56e6dc1b38cba0a610eea516b28fbff013b94f422b397",
361
+ "bytes": 368716
362
+ },
363
+ "v4": {
364
+ "path": "eval/heldout/v4/official/normalized/core.jsonl",
365
+ "sha256": "d6583f903e974d70747fae2944014828954dc4d7344ceddcbb2cd1b06fbfc799",
366
+ "bytes": 183671
367
+ },
368
+ "v5": {
369
+ "path": "eval/heldout/v5/official/normalized/core.jsonl",
370
+ "sha256": "f883164c4c83db8e42ebe03f7e165afbd2f81f2b541bcf8008a2cc06a60e1ee7",
371
+ "bytes": 200316
372
+ }
373
+ },
374
+ "transfer_rows": {
375
+ "path": "analysis/decision-benchmark-v3-official/scored/ROWS.json",
376
+ "sha256": "be29a7f3804ad84993ffb06f05ada7e2d8ed25317504c4fc17fe5dc306fd788c"
377
+ },
378
+ "evidence": [
379
+ {
380
+ "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
381
+ "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
382
+ },
383
+ {
384
+ "path": "analysis/decision-benchmark-v3-official/scored/REPORT.json",
385
+ "sha256": "2fdbaf7db1e4b877762e33a87b23a236a065eaa3078b14c2d7d579ad44f9ea24"
386
+ },
387
+ {
388
+ "path": "analysis/decision-benchmark-v3-official/INDEPENDENT-PROJECTION-REVIEW.json",
389
+ "sha256": "b9506539e2c4a863aef16f7dab0517c175e85863416bc622c60df05489a9eb52"
390
+ }
391
+ ]
392
+ },
393
+ "Kai": {
394
+ "panels": {
395
+ "old_core": {
396
+ "path": "results/public-roster-baselines-v1/Kai/worker/old_core/normalized.jsonl",
397
+ "sha256": "4758a99a75392b197c8902600a88edca65d6acea500e046e1ea8e4edd398a8a3"
398
+ },
399
+ "v3_core": {
400
+ "path": "results/public-roster-baselines-v1/Kai/worker/v3_core/normalized.jsonl",
401
+ "sha256": "7cdb285b88882a4c0d73c3da890056ef3b1d09e09bc01a893e73a12ff3bdbab8"
402
+ },
403
+ "v4": {
404
+ "path": "results/public-roster-baselines-v1/Kai/worker/v4/normalized.jsonl",
405
+ "sha256": "fa744ee2181fc4c379b748fdd4ee1080ffd877f43ebc9441d27a47013a2a9cec"
406
+ },
407
+ "v5": {
408
+ "path": "results/public-roster-baselines-v1/Kai/worker/v5/normalized.jsonl",
409
+ "sha256": "9d1e50d62733dbf06b824855938a9ee1311c030916cf173cc5b18fe066af48cd"
410
+ }
411
+ },
412
+ "transfer_rows": {
413
+ "path": "analysis/decision-benchmark-v3-baselines/Kai/ROWS.json",
414
+ "sha256": "854950cb3888a49a801982341576f4a2ebdacbdd7e7a7e32d16f5777e0cbda46"
415
+ },
416
+ "evidence": [
417
+ {
418
+ "path": "results/public-roster-baselines-v1/Kai/worker/COMPLETE.json",
419
+ "sha256": "39cc8bae7b59e02720123cd8efa9ba1939e02ef15a4bb3be89972b14eb7877cf"
420
+ },
421
+ {
422
+ "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
423
+ "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
424
+ },
425
+ {
426
+ "path": "analysis/decision-benchmark-v3-baselines/Kai/REPORT.json",
427
+ "sha256": "5ed1fd6a69d68fbd64bd45c4da9431a73fd6053574786ddd00323e0537fcefbc"
428
+ }
429
+ ]
430
+ },
431
+ "Laya-base": {
432
+ "panels": {
433
+ "v4": {
434
+ "path": "results/public-roster-baselines-v1/laya-english/worker/v4/normalized.jsonl",
435
+ "sha256": "31da1d4140da69d2677172707eb301843cf8fdfb43200cba93be910321f922ba"
436
+ },
437
+ "v5": {
438
+ "path": "results/public-roster-baselines-v1/laya-english/worker/v5/normalized.jsonl",
439
+ "sha256": "159614ceede814b578e1e025836a595c8c71a9bec4896caf6454d7bf54c4c44d"
440
+ },
441
+ "old_core": {
442
+ "path": "eval/heldout/primary-v2/laya-english/core/normalized.jsonl",
443
+ "sha256": "850643d2daeb47a6909985003736cab0f1f0e5a87631a35e27192f0733efd756"
444
+ },
445
+ "v3_core": {
446
+ "path": "eval/heldout/v3/baselines/laya-english/core/normalized.jsonl",
447
+ "sha256": "dec19caeb8074e25289e43a367392172da21f0c2ccdb403a059e2e4a2364a64b"
448
+ }
449
+ },
450
+ "transfer_rows": {
451
+ "path": "analysis/decision-benchmark-v3-baselines/laya-english/ROWS.json",
452
+ "sha256": "19138909259400facb5d3d5402d0d736f870695fe3ee46019f2e438088c57275"
453
+ },
454
+ "evidence": [
455
+ {
456
+ "path": "results/public-roster-baselines-v1/laya-english/worker/COMPLETE.json",
457
+ "sha256": "cef0da797ae14cebaf1e4b9edf9b16c3cd9a412a30435d7d2517b8ba6dbe8b98"
458
+ },
459
+ {
460
+ "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
461
+ "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
462
+ },
463
+ {
464
+ "path": "analysis/decision-benchmark-v3-baselines/laya-english/REPORT.json",
465
+ "sha256": "cb5935bb399dcf23ffe24f855452b2e074bcbb3d023cfc7dd439342e54f02c6b"
466
+ },
467
+ {
468
+ "path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
469
+ "sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
470
+ }
471
+ ]
472
+ },
473
+ "Laya-multilingual": {
474
+ "panels": {
475
+ "v4": {
476
+ "path": "results/public-roster-baselines-v1/laya-multilingual/worker/v4/normalized.jsonl",
477
+ "sha256": "814c1aa1f4d400942984ecb3786d3fd5fd15286d9b767a519c91993577795290"
478
+ },
479
+ "v5": {
480
+ "path": "results/public-roster-baselines-v1/laya-multilingual/worker/v5/normalized.jsonl",
481
+ "sha256": "6d4dbc16c7b977b15895e0e313dabc6d9db43df75ec56e89373d91f4f13ec676"
482
+ },
483
+ "old_core": {
484
+ "path": "eval/heldout/primary-v2/laya-multilingual/core/normalized.jsonl",
485
+ "sha256": "9ea39841c4d5901cd9ab860c980d2f3b55358be289c7e01b9b47982927bf849d"
486
+ },
487
+ "v3_core": {
488
+ "path": "eval/heldout/v3/baselines/laya-multilingual/core/normalized.jsonl",
489
+ "sha256": "a0b7ce9daf4be2c7831bcab6b5e96c8dcb6bef133d39475f72241ab18faccdd2"
490
+ }
491
+ },
492
+ "transfer_rows": {
493
+ "path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/ROWS.json",
494
+ "sha256": "d32db6cf6576f287210d31336456497f9d6926179366db29304659baf9783b43"
495
+ },
496
+ "evidence": [
497
+ {
498
+ "path": "results/public-roster-baselines-v1/laya-multilingual/worker/COMPLETE.json",
499
+ "sha256": "b190d683fc4c4762c2805a03d52628a53e125d228a3105c8ac2ddda00a5364e4"
500
+ },
501
+ {
502
+ "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
503
+ "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
504
+ },
505
+ {
506
+ "path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/REPORT.json",
507
+ "sha256": "d20d3d58f76ca2758545f83d358abd747f86d581c897d1fb521371d1000d6154"
508
+ },
509
+ {
510
+ "path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
511
+ "sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
512
+ }
513
+ ]
514
+ },
515
+ "Qwen3.5-9B": {
516
+ "panels": {
517
+ "old_core": {
518
+ "path": "results/lux-new-host-v1/quality/base9b/old_core/normalized.jsonl",
519
+ "sha256": "16ccd1c30bbf29560a7ce1a288143692c5d2ed69cc7afebdc146d5bcc4b77611"
520
+ },
521
+ "v3_core": {
522
+ "path": "results/lux-new-host-v1/quality/base9b/v3_core/normalized.jsonl",
523
+ "sha256": "d081a8158863e85d5e49cb853fee87192119767124f53f9e6c782392f8edb018"
524
+ },
525
+ "v4": {
526
+ "path": "results/lux-new-host-v1/quality/base9b/v4/normalized.jsonl",
527
+ "sha256": "cc2775884d84c49e5155afbebc7f4562ec86f6f6d32765ffcbdafe653ca7d132"
528
+ },
529
+ "v5": {
530
+ "path": "results/lux-new-host-v1/quality/base9b/v5/normalized.jsonl",
531
+ "sha256": "ce6bf5c286e5920055d7de36fbd910311f7d3185380c63681c7baea6d0e3494a"
532
+ }
533
+ },
534
+ "transfer_rows": {
535
+ "path": "analysis/decision-benchmark-v3-baselines/Qwen3.5-9B/ROWS.json",
536
+ "sha256": "5acdd41f616a3ba63a5407af9c8371268012792fd324b70230d899bbb37101b4"
537
+ },
538
+ "evidence": [
539
+ {
540
+ "path": "results/lux-new-host-v1/quality/base9b/COMPLETE.json",
541
+ "sha256": "702c65f5c290f5cdda0495479f361101dbe976330233c5359455ba736245c3aa"
542
+ },
543
+ {
544
+ "path": "results/lux-new-host-v1/quality/base9b/metadata.json",
545
+ "sha256": "eaa29f71aba60e86068e6c5d1b5782ea73b0214b81a297b57419c4f0a2fc6c04"
546
+ },
547
+ {
548
+ "path": "results/lux-new-host-v1/quality/base9b/old_core/complete.json",
549
+ "sha256": "d28f2a646a908fc775e3727569a608f30ae5820f88a6173750a1031604e84f47"
550
+ },
551
+ {
552
+ "path": "results/lux-new-host-v1/quality/base9b/v3_core/complete.json",
553
+ "sha256": "3a10bd5e2121a81f8ec6f76b2526723a17689ec34d13f062c5259983203992cc"
554
+ },
555
+ {
556
+ "path": "results/lux-new-host-v1/quality/base9b/v4/complete.json",
557
+ "sha256": "073239ab7dd21768bccdfa4aaf3217c99dc7ff9b00fc308ffbe889c72fe002b4"
558
+ },
559
+ {
560
+ "path": "results/lux-new-host-v1/quality/base9b/v5/complete.json",
561
+ "sha256": "9a8d1e8e8db066f1f6594f137a7a1571dd5184aae363a9f919e44b634b782bba"
562
+ },
563
+ {
564
+ "path": "analysis/decision-benchmark-v3-baselines/Qwen3.5-9B/REPORT.json",
565
+ "sha256": "31b0a92bbbe7a0d259cb8e296a10b04b2bdd685b394eabdc5fbf6da1556ff12d"
566
+ },
567
+ {
568
+ "path": "analysis/decision-benchmark-v3-baselines/Qwen3.5-9B-projected/COMPLETE.json",
569
+ "sha256": "eaff79adb3a6a6bad61253a01b4c4e22cfa3d9969cd963cd82a3685755b77ec1"
570
+ }
571
+ ]
572
+ },
573
+ "Lux": {
574
+ "panels": {
575
+ "old_core": {
576
+ "path": "results/lux-new-host-v1/quality/lux/old_core/normalized.jsonl",
577
+ "sha256": "a659a7c7fef2871bbf806e842ea8ca81f645dcf610f0f2f6622d8d32e76a0836"
578
+ },
579
+ "v3_core": {
580
+ "path": "results/lux-new-host-v1/quality/lux/v3_core/normalized.jsonl",
581
+ "sha256": "da8903a07e36a9687f6ee1d0fd8dbc684e61c49ef5fea30c346e46bff5a90e94"
582
+ },
583
+ "v4": {
584
+ "path": "results/lux-new-host-v1/quality/lux/v4/normalized.jsonl",
585
+ "sha256": "26cd7482693ec1ffa1a93a156abfc88378e4872e320f3b9b775763635086d488"
586
+ },
587
+ "v5": {
588
+ "path": "results/lux-new-host-v1/quality/lux/v5/normalized.jsonl",
589
+ "sha256": "759c773e629a16317271de850780a8734533c795046122f850914c4482d1e72b"
590
+ }
591
+ },
592
+ "transfer_rows": {
593
+ "path": "analysis/decision-benchmark-v3-baselines/Lux/ROWS.json",
594
+ "sha256": "bafb55fc86762061e537382360eda14a8a8f8c5e4e111c2fda1f82cd6c06793f"
595
+ },
596
+ "evidence": [
597
+ {
598
+ "path": "results/lux-new-host-v1/quality/lux/COMPLETE.json",
599
+ "sha256": "9600ff792cbda13af5505c804cf63af94d2ee895f2912f0645a7d493ae1b7a9c"
600
+ },
601
+ {
602
+ "path": "results/lux-new-host-v1/quality/lux/metadata.json",
603
+ "sha256": "0d4509fae7eb3630118eaebc24e75a7b02f717d9a0ad74811fa27358c33c9cc5"
604
+ },
605
+ {
606
+ "path": "analysis/decision-benchmark-v3-baselines/Lux-projected/COMPLETE.json",
607
+ "sha256": "f58d02df78f297d4feebe39facfac072951111d8d6a19e2c1069f1b90250c622"
608
+ },
609
+ {
610
+ "path": "analysis/decision-benchmark-v3-baselines/Lux/REPORT.json",
611
+ "sha256": "f31703c342efe043a1e1534da189808bb7e0de0641fbaf3f91faa4a1f928a299"
612
+ },
613
+ {
614
+ "path": "analysis/decoder4b/lux9b-training-v1/CHECKPOINT-HELDOUT-COLLECTED.json",
615
+ "sha256": "d1c42515ab225b27ea011a6cf2262fe12dd7001e3b4d2cb64a23e7642c7a0d91"
616
+ }
617
+ ]
618
+ },
619
+ "Decider": {
620
+ "panels": {
621
+ "old_core": {
622
+ "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/decider/quality/normalized.jsonl",
623
+ "sha256": "e1b14f0f9ec521bffbde5945684bc2410a413b6b6418d65ac4fc702a7b5e7937"
624
+ },
625
+ "v3_core": {
626
+ "path": "eval/heldout/v3/baselines/decider/core/normalized.jsonl",
627
+ "sha256": "12a1f554decf1aff608743e7b4a44681289ea084ae25f383e519b1f34917d46f"
628
+ },
629
+ "v4": {
630
+ "path": "eval/heldout/v4/baselines/decider/core/normalized.jsonl",
631
+ "sha256": "626ef366ae471ec10dfb89ef2ff2f7b29a6c879976d3ff14fcbb2bafd7d9b045"
632
+ },
633
+ "v5": {
634
+ "path": "eval/heldout/v5/open-baselines-v1/decider/core/normalized.jsonl",
635
+ "sha256": "ae16b17ac8e4208940c0c3e9e04f25e94686b5d16bf46f64f53669baf41aeb71"
636
+ }
637
+ },
638
+ "transfer_rows": {
639
+ "path": "analysis/decision-benchmark-v3-baselines/decider/ROWS.json",
640
+ "sha256": "6da5b5692a8a0cea3dc8f6d02d911fafd9bbf62e286a9be8db27c89f77a897e4"
641
+ },
642
+ "evidence": [
643
+ {
644
+ "path": "results/accelerated-transfer-baselines-v1/decider/worker/COMPLETE.json",
645
+ "sha256": "e1ee0df8df59ec3eb2b49ba947a8508994a24bce226b7f2cccbe25f6835b99af"
646
+ },
647
+ {
648
+ "path": "results/accelerated-transfer-baselines-v1/decider/worker/METADATA.json",
649
+ "sha256": "f00061c1b0bb642bf1353aae867d19df988c01b0428b58206b31249521bcbff8"
650
+ },
651
+ {
652
+ "path": "results/accelerated-transfer-baselines-v1/decider/worker/predictions.jsonl",
653
+ "sha256": "f69e87fa65b09289e6861c1d789270aea5a92d23ac47817342647ed08f406df7"
654
+ },
655
+ {
656
+ "path": "analysis/decision-benchmark-v3-baselines/decider/REPORT.json",
657
+ "sha256": "5aae029f34524dfebc8e79128da8c39229c5747461dfce07ca59956da4416951"
658
+ },
659
+ {
660
+ "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
661
+ "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
662
+ },
663
+ {
664
+ "path": "analysis/decision-benchmark-v3-baselines/decider/COMPOSABLE-MANIFEST.json",
665
+ "sha256": "a49848e86feafaa4f36a33287910035e7605e344791aae139305b78cb304a77f"
666
+ }
667
+ ]
668
+ },
669
+ "Qwen3.5-2B": {
670
+ "panels": {
671
+ "old_core": {
672
+ "path": "eval/heldout/primary-v2/base-2b/core/normalized.jsonl",
673
+ "sha256": "c11a94d8f833d83427ea984a4f069aefaa3fa10a2b1085fb3a03542727f819cf"
674
+ },
675
+ "v3_core": {
676
+ "path": "eval/heldout/v3/baselines/base-2b/core/normalized.jsonl",
677
+ "sha256": "cc51471716342ecd466b50d5dd9b4ba805f5a6873cd96075b56d7c2e8336175e"
678
+ },
679
+ "v4": {
680
+ "path": "eval/heldout/v4/baselines/base-2b/core/normalized.jsonl",
681
+ "sha256": "93d84cdbe607691e2cef35ece3aa2dd5900f6410b0d72796b86d74c9156f3cdd"
682
+ },
683
+ "v5": {
684
+ "path": "eval/heldout/v5/open-baselines-v1/base-2b/core/normalized.jsonl",
685
+ "sha256": "75740c4279eb8dcc9c90b55f155916b3748c887e8c9d9e1d3a47edf7f4dbc34b"
686
+ }
687
+ },
688
+ "transfer_rows": {
689
+ "path": "analysis/decision-benchmark-v3-baselines/base-2b/ROWS.json",
690
+ "sha256": "62bf4f977d03eabd4f051aca5c8f7204f350eb995545aedeefa7bfa00bcbd115"
691
+ },
692
+ "evidence": [
693
+ {
694
+ "path": "results/accelerated-transfer-baselines-v1/base-2b/worker/COMPLETE.json",
695
+ "sha256": "0353d786b760d491c2a5ff3f8363ecd27a537f25a42dbeb2cea4db5d5eb77910"
696
+ },
697
+ {
698
+ "path": "results/accelerated-transfer-baselines-v1/base-2b/worker/METADATA.json",
699
+ "sha256": "745151846ade07431bc8ac964894bbd6ed234fb6a4665fcce223d91e09acaeeb"
700
+ },
701
+ {
702
+ "path": "results/accelerated-transfer-baselines-v1/base-2b/worker/predictions.jsonl",
703
+ "sha256": "347fc785ee59d5ca7694357b68a6c2e7bc7f863f1e26d94fc9b989391ed7d5fe"
704
+ },
705
+ {
706
+ "path": "analysis/decision-benchmark-v3-baselines/base-2b/REPORT.json",
707
+ "sha256": "9eb1db12e826a75dcda322960478bb153c4c1c24e5f5853d29a761b82bea36e7"
708
+ },
709
+ {
710
+ "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
711
+ "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
712
+ },
713
+ {
714
+ "path": "analysis/decision-benchmark-v3-baselines/base-2b/COMPOSABLE-MANIFEST.json",
715
+ "sha256": "c3e6f28b600776686afc56c1386cabf0869e69d5ca79f3ae6a0b17189eee968e"
716
+ }
717
+ ]
718
+ },
719
+ "Qwen3.5-4B": {
720
+ "panels": {
721
+ "old_core": {
722
+ "path": "eval/heldout/primary-v2/base-4b/core/normalized.jsonl",
723
+ "sha256": "d20ad845066ff81caaf60ebee44df941872bdc642da78b5f3b3da4a33dd9a08d"
724
+ },
725
+ "v3_core": {
726
+ "path": "eval/heldout/v3/baselines/base-4b/core/normalized.jsonl",
727
+ "sha256": "7541893d2b2ec0ec2ffc513f50c568eaef860b65bcef9a1747996c116fff9661"
728
+ },
729
+ "v4": {
730
+ "path": "eval/heldout/v4/baselines/base-4b/core/normalized.jsonl",
731
+ "sha256": "7a6d252ccaa72bf54b94445190bbad023fb465a3eaf4cf657532e2cdd4f750f5"
732
+ },
733
+ "v5": {
734
+ "path": "eval/heldout/v5/open-baselines-v1/base-4b/core/normalized.jsonl",
735
+ "sha256": "c01ad5daed65267e9aec40ed9e97ece1b6511e4cab17ebff09339b41524fb7f1"
736
+ }
737
+ },
738
+ "transfer_rows": {
739
+ "path": "analysis/decision-benchmark-v3-baselines/base-4b/ROWS.json",
740
+ "sha256": "37ab17d94d640ee46283e1f23723b9ea05133c22a8a9b62a97c756a89c5de785"
741
+ },
742
+ "evidence": [
743
+ {
744
+ "path": "results/remaining-transfer-baselines-v1/base-4b/worker/COMPLETE.json",
745
+ "sha256": "72abc4f87a9234f95eb3b0a72a9e7d66798779d5982c757faffb610003b1b242"
746
+ },
747
+ {
748
+ "path": "results/remaining-transfer-baselines-v1/base-4b/worker/METADATA.json",
749
+ "sha256": "871fed92cf762187015e59174a042808dbf3fc88f7791aa0ddd530ae923196ab"
750
+ },
751
+ {
752
+ "path": "results/remaining-transfer-baselines-v1/base-4b/worker/predictions.jsonl",
753
+ "sha256": "8c47edda469f8914eccc2ecafa91c245752812e830cb9f29445653102522854a"
754
+ },
755
+ {
756
+ "path": "analysis/decision-benchmark-v3-baselines/base-4b/REPORT.json",
757
+ "sha256": "c2a0aafbe2b7c85b3c07a4b9843182a3361654f57883311155aa3b05cfc7002d"
758
+ },
759
+ {
760
+ "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
761
+ "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
762
+ },
763
+ {
764
+ "path": "analysis/decision-benchmark-v3-baselines/base-4b/COMPOSABLE-MANIFEST.json",
765
+ "sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
766
+ }
767
+ ]
768
+ },
769
+ "llm2jev-2b": {
770
+ "panels": {
771
+ "old_core": {
772
+ "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard0/llm2jev-2b/quality/normalized.jsonl",
773
+ "sha256": "5f0aaf2be47d11c42ecb228aeb6f317bf2074c80d398a66969f06490d5847760"
774
+ },
775
+ "v3_core": {
776
+ "path": "eval/heldout/v3/baselines/llm2jev-2b/core/normalized.jsonl",
777
+ "sha256": "339cd97706404b3d1e5af9d07ae6d3efc79cbb40310cfcf1406c64477e24e460"
778
+ },
779
+ "v4": {
780
+ "path": "results/open-reference-gap-v1/llm2jev-2b/v4/normalized.jsonl",
781
+ "sha256": "98e6b3bc2113e18e00442a86b085e06981560f43cf6a8807c3d1818b3fd27caf"
782
+ },
783
+ "v5": {
784
+ "path": "results/open-reference-gap-v1/llm2jev-2b/v5/normalized.jsonl",
785
+ "sha256": "36b9d9e96bae101eb3db6600c412585f48551ff5908072963a5dd435632e4d65"
786
+ }
787
+ },
788
+ "transfer_rows": {
789
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/ROWS.json",
790
+ "sha256": "a37a6b19549896854fd0fb9bb32f594bb8cb60b840aa49ead0c004dcea00778e"
791
+ },
792
+ "evidence": [
793
+ {
794
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/COMPLETE.json",
795
+ "sha256": "ccd7e1f0d6a4193c629046961eb7171b1278a8dbae015ff216396ceac1f008ef"
796
+ },
797
+ {
798
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/METADATA.json",
799
+ "sha256": "b0d0aa9840976e0330e676fdd67ce90230dbfec7502452956260188924583b2b"
800
+ },
801
+ {
802
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/predictions.jsonl",
803
+ "sha256": "67a217eea9af720846f653d780973a64fff93111d8dcb9fa09b8b6f9594328a4"
804
+ },
805
+ {
806
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/REPORT.json",
807
+ "sha256": "bb2bf635114fc805e6fef66dce70ac6ceca126008b3e8365871a6d12d9cd6e09"
808
+ },
809
+ {
810
+ "path": "eval/open-reference-comparison-v1/REGISTRY.json",
811
+ "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
812
+ },
813
+ {
814
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/COMPOSABLE-MANIFEST.json",
815
+ "sha256": "071d3b47bd6a81e256e11be2ed2965c7416244bd512bd35be8baef5b8ebcb2be"
816
+ }
817
+ ]
818
+ },
819
+ "llm2jev-4b": {
820
+ "panels": {
821
+ "old_core": {
822
+ "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/llm2jev-4b/quality/normalized.jsonl",
823
+ "sha256": "8545cdc4ab9e9f870fdce1c1859632b8bac7505f6e961b4cc07e4f3b816328ce"
824
+ },
825
+ "v3_core": {
826
+ "path": "eval/heldout/v3/baselines/llm2jev-4b/core/normalized.jsonl",
827
+ "sha256": "b3bd38c90daf0bc3d77fea9a3ed0612049b07865bde5102356f82f7456002c97"
828
+ },
829
+ "v4": {
830
+ "path": "results/open-reference-gap-v1/llm2jev-4b/v4/normalized.jsonl",
831
+ "sha256": "409f8d2c57a1f8a4ef7dd18f3adb6cb8ee2ed544566206ccc909d54236ca95b7"
832
+ },
833
+ "v5": {
834
+ "path": "results/open-reference-gap-v1/llm2jev-4b/v5/normalized.jsonl",
835
+ "sha256": "d458e7e11cd2b1d66c7db6044408042630a48a474c9d3997141fc85dfb427f24"
836
+ }
837
+ },
838
+ "transfer_rows": {
839
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/ROWS.json",
840
+ "sha256": "abc8f7c04170544f5c4bef7a9af70ce0d7a25bfc4b4b8b1ba3e552413273bb7d"
841
+ },
842
+ "evidence": [
843
+ {
844
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/COMPLETE.json",
845
+ "sha256": "d7745d5ec1e08100629c55e2c7518ff408462e6adea4ebfb01b7b401b212c4b6"
846
+ },
847
+ {
848
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/METADATA.json",
849
+ "sha256": "6a522387522b92a0c01cbd8f32cdaa154289318a98564297d2f5723bc4c3dee0"
850
+ },
851
+ {
852
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/predictions.jsonl",
853
+ "sha256": "61bea4eb25014f6277cf81010872c1724e6d8f07a89803063649a21160b02edc"
854
+ },
855
+ {
856
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/REPORT.json",
857
+ "sha256": "1097d4eeb4abc792cac569e4793c469a332dc24bbdf5568a242ce4cee4c03759"
858
+ },
859
+ {
860
+ "path": "eval/open-reference-comparison-v1/REGISTRY.json",
861
+ "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
862
+ },
863
+ {
864
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/COMPOSABLE-MANIFEST.json",
865
+ "sha256": "82a7589e866610251b5b52f5adb20ab5c7ed9fa6fb2a184c9d6a3367576171e7"
866
+ }
867
+ ]
868
+ },
869
+ "nimble-9b": {
870
+ "panels": {
871
+ "old_core": {
872
+ "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/nimble-9b/quality/normalized.jsonl",
873
+ "sha256": "6fefc71c6fbea1106894c1081057c2c7a253fa4ebc2e6aecfe23abc347be08bb"
874
+ },
875
+ "v3_core": {
876
+ "path": "eval/heldout/v3/baselines/nimble-9b/core/normalized.jsonl",
877
+ "sha256": "ab5323762fc5d34961c0fbce78497627528b696c60e18b880d671b468db97bf9"
878
+ },
879
+ "v4": {
880
+ "path": "results/open-reference-gap-v1/nimble-9b/v4/normalized.jsonl",
881
+ "sha256": "6ae263cef9cafdbcd35cd1f1dcc5b1bd920d3e6023d4fc487970286c04ec8d05"
882
+ },
883
+ "v5": {
884
+ "path": "eval/heldout/v5/open-baselines-v1/nimble-9b/core/normalized.jsonl",
885
+ "sha256": "0fc6dae372265ef62ea87ce9dc07b63a77ccd33e3aa48e15754f5ab98c32580a"
886
+ }
887
+ },
888
+ "transfer_rows": {
889
+ "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/ROWS.json",
890
+ "sha256": "f55f29f973b0974adc7fc3fd582e2e5c0414f440e4a094a7ec8e152685e62a28"
891
+ },
892
+ "evidence": [
893
+ {
894
+ "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/COMPLETE.json",
895
+ "sha256": "efc5a1be536227db99ff9937b84e9a5bc688fc497c0e4a84d8261144cea8a51d"
896
+ },
897
+ {
898
+ "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/METADATA.json",
899
+ "sha256": "77eaa373ec6a74b483ac9b0cda739ecd8cce95a046266cd7a28488a768b82940"
900
+ },
901
+ {
902
+ "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/predictions.jsonl",
903
+ "sha256": "ad2212fa691c662c7917b33315439844339252778fdaefcc3080cdf46e00a434"
904
+ },
905
+ {
906
+ "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/REPORT.json",
907
+ "sha256": "b6065189b7522e5d0d2b7b088b673ba26212bee5fda2f0d7c6cd67fc9cc93136"
908
+ },
909
+ {
910
+ "path": "eval/open-reference-comparison-v1/REGISTRY.json",
911
+ "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
912
+ },
913
+ {
914
+ "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/COMPOSABLE-MANIFEST.json",
915
+ "sha256": "7423a2430984cf8e7eed2c5933ceeb0e2894d252e30e447d1294b6b6b08192bd"
916
+ }
917
+ ]
918
+ }
919
+ },
920
+ "task_rows": 54,
921
+ "rank_public_models": [
922
+ "Jev",
923
+ "Lux",
924
+ "Nox",
925
+ "kev-9b",
926
+ "kev-4b",
927
+ "Qwen3.5-9B",
928
+ "Decider",
929
+ "Qwen3.5-4B",
930
+ "Sol",
931
+ "kev-0.8b",
932
+ "Qwen3.5-2B",
933
+ "Laya-base",
934
+ "Laya-multilingual",
935
+ "Kai"
936
+ ],
937
+ "internal_extra_models_not_public_rank": [
938
+ "llm2jev-2b",
939
+ "llm2jev-4b",
940
+ "nimble-9b"
941
+ ],
942
+ "laya_parameter_audit": {
943
+ "utc": "2026-09-22T05:01:58.295720+00:00",
944
+ "torch": "2.12.0+git6bbd260",
945
+ "transformers": "5.17.0",
946
+ "device": "meta, CPU-only process; no GPU devices exposed",
947
+ "models": {
948
+ "laya-english": {
949
+ "serialized_scalar_count": 421293830,
950
+ "unique_parameter_count": 421293827,
951
+ "parameter_entries": 205,
952
+ "named_parameter_entries_with_duplicates": 205,
953
+ "tied_parameter_aliases": [],
954
+ "persistent_buffers": {
955
+ "temperature": [
956
+ 3
957
+ ]
958
+ },
959
+ "persistent_buffer_scalar_count": 3,
960
+ "nonpersistent_buffers": {
961
+ "encoder.rotary_emb.full_attention_inv_freq": [
962
+ 32
963
+ ],
964
+ "encoder.rotary_emb.full_attention_original_inv_freq": [
965
+ 32
966
+ ],
967
+ "encoder.rotary_emb.sliding_attention_inv_freq": [
968
+ 32
969
+ ],
970
+ "encoder.rotary_emb.sliding_attention_original_inv_freq": [
971
+ 32
972
+ ]
973
+ },
974
+ "exact_state_dict_shapes_match_header": true,
975
+ "header_sha256": "3a42d8e7a96d5aa22c5bc9475d7ff092e32a832ae7c6973688f66f7a84048c04",
976
+ "source_sha256": "8d83611d480c971d640a7b7d3aa2f2219c5e8455e9cc2329fd073681bd8be23e",
977
+ "encoder_config_sha256": "bf3ab80598fdccf414855a2ce80f22859e4492d06ca8a62ddd1cfb63972f8979",
978
+ "agent_config_sha256": "ae287b56bbcf5f8c4f4541ae9dfd00c914c4c48b940b8398c3058af37ba92bbd",
979
+ "architecture": "ModernBertModel",
980
+ "interpretation": "Complete decision model, including encoder and heads; temperature calibration buffer excluded. No tied parameter aliases in the instantiated published model."
981
+ },
982
+ "laya-multilingual": {
983
+ "serialized_scalar_count": 321908998,
984
+ "unique_parameter_count": 321908995,
985
+ "parameter_entries": 169,
986
+ "named_parameter_entries_with_duplicates": 169,
987
+ "tied_parameter_aliases": [],
988
+ "persistent_buffers": {
989
+ "temperature": [
990
+ 3
991
+ ]
992
+ },
993
+ "persistent_buffer_scalar_count": 3,
994
+ "nonpersistent_buffers": {
995
+ "encoder.rotary_emb.full_attention_inv_freq": [
996
+ 32
997
+ ],
998
+ "encoder.rotary_emb.full_attention_original_inv_freq": [
999
+ 32
1000
+ ],
1001
+ "encoder.rotary_emb.sliding_attention_inv_freq": [
1002
+ 32
1003
+ ],
1004
+ "encoder.rotary_emb.sliding_attention_original_inv_freq": [
1005
+ 32
1006
+ ]
1007
+ },
1008
+ "exact_state_dict_shapes_match_header": true,
1009
+ "header_sha256": "22ab94329063133fdd2944997b906bb6076d2cfd3e43eb05d7ea92187e9c3984",
1010
+ "source_sha256": "8d83611d480c971d640a7b7d3aa2f2219c5e8455e9cc2329fd073681bd8be23e",
1011
+ "encoder_config_sha256": "83f6916d13ef0f556ac461f28308dc2bffa7ebeadee8ec9e2db5812020ea5bb4",
1012
+ "agent_config_sha256": "25061739243b617ad88d1219ba6f8a9c86c5881ca28df024fa2d9b3b2fcc30c6",
1013
+ "architecture": "ModernBertModel",
1014
+ "interpretation": "Complete decision model, including encoder and heads; temperature calibration buffer excluded. No tied parameter aliases in the instantiated published model."
1015
+ }
1016
+ },
1017
+ "source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
1018
+ }
1019
+ }
metrics/question-scaling.json CHANGED
@@ -1,225 +1,16 @@
1
  {
2
- "all_types": {
3
- "Nox-retention-selected": {
4
- "choice-q1": {
5
- "samples": 30,
6
- "p50_ms": 32.731553500000004,
7
- "p95_ms": 33.4735869,
8
- "requests_per_second_from_median": 30.551559369157346,
9
- "questions_per_second_from_median": 30.551559369157346,
10
- "measured_requests_per_second": 30.5029057383203,
11
- "measured_questions_per_second": 30.5029057383203,
12
- "max_allocated_bytes": 8649496576,
13
- "max_reserved_bytes": 8923381760,
14
- "input_tokens": 309
15
- },
16
- "choice-q2": {
17
- "samples": 30,
18
- "p50_ms": 33.3187065,
19
- "p95_ms": 33.56022575,
20
- "requests_per_second_from_median": 30.01316992903071,
21
- "questions_per_second_from_median": 60.02633985806142,
22
- "measured_requests_per_second": 30.051085192016746,
23
- "measured_questions_per_second": 60.10217038403349,
24
- "max_allocated_bytes": 8694344704,
25
- "max_reserved_bytes": 8938061824,
26
- "input_tokens": 618
27
- },
28
- "choice-q4": {
29
- "samples": 30,
30
- "p50_ms": 41.779055,
31
- "p95_ms": 42.214628999999995,
32
- "requests_per_second_from_median": 23.935438463124644,
33
- "questions_per_second_from_median": 95.74175385249858,
34
- "measured_requests_per_second": 23.884399734574895,
35
- "measured_questions_per_second": 95.53759893829958,
36
- "max_allocated_bytes": 8783385600,
37
- "max_reserved_bytes": 9141485568,
38
- "input_tokens": 1236
39
- },
40
- "choice-q8": {
41
- "samples": 30,
42
- "p50_ms": 69.641899,
43
- "p95_ms": 70.72795845,
44
- "requests_per_second_from_median": 14.359171911725154,
45
- "questions_per_second_from_median": 114.87337529380123,
46
- "measured_requests_per_second": 14.311513624064359,
47
- "measured_questions_per_second": 114.49210899251487,
48
- "max_allocated_bytes": 8963302400,
49
- "max_reserved_bytes": 9376366592,
50
- "input_tokens": 2472
51
- },
52
- "choice-q16": {
53
- "samples": 30,
54
- "p50_ms": 141.0197655,
55
- "p95_ms": 141.85338695,
56
- "requests_per_second_from_median": 7.091204530474134,
57
- "questions_per_second_from_median": 113.45927248758615,
58
- "measured_requests_per_second": 7.1075954042236935,
59
- "measured_questions_per_second": 113.7215264675791,
60
- "max_allocated_bytes": 8962254336,
61
- "max_reserved_bytes": 9244246016,
62
- "input_tokens": 4944
63
- },
64
- "choice-q32": {
65
- "samples": 30,
66
- "p50_ms": 282.0275835,
67
- "p95_ms": 283.68355145,
68
- "requests_per_second_from_median": 3.545752467151852,
69
- "questions_per_second_from_median": 113.46407894885927,
70
- "measured_requests_per_second": 3.5466119610851,
71
- "measured_questions_per_second": 113.4915827547232,
72
- "max_allocated_bytes": 8963302912,
73
- "max_reserved_bytes": 9141485568,
74
- "input_tokens": 9888
75
- },
76
- "noul-q1": {
77
- "samples": 30,
78
- "p50_ms": 32.685345,
79
- "p95_ms": 33.000833899999996,
80
- "requests_per_second_from_median": 30.594751256258732,
81
- "questions_per_second_from_median": 30.594751256258732,
82
- "measured_requests_per_second": 30.61657989325799,
83
- "measured_questions_per_second": 30.61657989325799,
84
- "max_allocated_bytes": 8645022720,
85
- "max_reserved_bytes": 8917090304,
86
- "input_tokens": 262
87
- },
88
- "noul-q2": {
89
- "samples": 30,
90
- "p50_ms": 32.822731000000005,
91
- "p95_ms": 33.580505,
92
- "requests_per_second_from_median": 30.466690903934833,
93
- "questions_per_second_from_median": 60.933381807869665,
94
- "measured_requests_per_second": 30.306271635312687,
95
- "measured_questions_per_second": 60.61254327062537,
96
- "max_allocated_bytes": 8686444544,
97
- "max_reserved_bytes": 9216983040,
98
- "input_tokens": 524
99
- },
100
- "noul-q4": {
101
- "samples": 30,
102
- "p50_ms": 38.8360825,
103
- "p95_ms": 39.62854105,
104
- "requests_per_second_from_median": 25.749250069185013,
105
- "questions_per_second_from_median": 102.99700027674005,
106
- "measured_requests_per_second": 25.627138601267134,
107
- "measured_questions_per_second": 102.50855440506854,
108
- "max_allocated_bytes": 8767323136,
109
- "max_reserved_bytes": 9191817216,
110
- "input_tokens": 1048
111
- },
112
- "noul-q8": {
113
- "samples": 30,
114
- "p50_ms": 64.278516,
115
- "p95_ms": 65.07748314999999,
116
- "requests_per_second_from_median": 15.557297558020787,
117
- "questions_per_second_from_median": 124.4583804641663,
118
- "measured_requests_per_second": 15.523243079930335,
119
- "measured_questions_per_second": 124.18594463944268,
120
- "max_allocated_bytes": 8932226048,
121
- "max_reserved_bytes": 9141485568,
122
- "input_tokens": 2096
123
- },
124
- "noul-q16": {
125
- "samples": 30,
126
- "p50_ms": 129.2654275,
127
- "p95_ms": 130.59987905,
128
- "requests_per_second_from_median": 7.7360205225794045,
129
- "questions_per_second_from_median": 123.77632836127047,
130
- "measured_requests_per_second": 7.725552951411574,
131
- "measured_questions_per_second": 123.60884722258518,
132
- "max_allocated_bytes": 8931440128,
133
- "max_reserved_bytes": 9244246016,
134
- "input_tokens": 4192
135
- },
136
- "noul-q32": {
137
- "samples": 30,
138
- "p50_ms": 259.156278,
139
- "p95_ms": 261.59688335,
140
- "requests_per_second_from_median": 3.8586755748977075,
141
- "questions_per_second_from_median": 123.47761839672664,
142
- "measured_requests_per_second": 3.8613862674541757,
143
- "measured_questions_per_second": 123.56436055853362,
144
- "max_allocated_bytes": 8932226560,
145
- "max_reserved_bytes": 9244246016,
146
- "input_tokens": 8384
147
- },
148
- "score-q1": {
149
- "samples": 30,
150
- "p50_ms": 33.319307,
151
- "p95_ms": 34.269436,
152
- "requests_per_second_from_median": 30.01262901416287,
153
- "questions_per_second_from_median": 30.01262901416287,
154
- "measured_requests_per_second": 29.995543652066246,
155
- "measured_questions_per_second": 29.995543652066246,
156
- "max_allocated_bytes": 8648972288,
157
- "max_reserved_bytes": 9244246016,
158
- "input_tokens": 307
159
- },
160
- "score-q2": {
161
- "samples": 30,
162
- "p50_ms": 33.3555535,
163
- "p95_ms": 33.76658145,
164
- "requests_per_second_from_median": 29.980015171986278,
165
- "questions_per_second_from_median": 59.960030343972555,
166
- "measured_requests_per_second": 30.07546038274141,
167
- "measured_questions_per_second": 60.15092076548282,
168
- "max_allocated_bytes": 8693689344,
169
- "max_reserved_bytes": 9141485568,
170
- "input_tokens": 614
171
- },
172
- "score-q4": {
173
- "samples": 30,
174
- "p50_ms": 41.8594245,
175
- "p95_ms": 42.2104965,
176
- "requests_per_second_from_median": 23.889482761522437,
177
- "questions_per_second_from_median": 95.55793104608975,
178
- "measured_requests_per_second": 23.829400385190663,
179
- "measured_questions_per_second": 95.31760154076265,
180
- "max_allocated_bytes": 8782337024,
181
- "max_reserved_bytes": 9244246016,
182
- "input_tokens": 1228
183
- },
184
- "score-q8": {
185
- "samples": 30,
186
- "p50_ms": 69.734315,
187
- "p95_ms": 70.12355445,
188
- "requests_per_second_from_median": 14.340142295797987,
189
- "questions_per_second_from_median": 114.7211383663839,
190
- "measured_requests_per_second": 14.393887061255466,
191
- "measured_questions_per_second": 115.15109649004373,
192
- "max_allocated_bytes": 8963302400,
193
- "max_reserved_bytes": 9376366592,
194
- "input_tokens": 2456
195
- },
196
- "score-q16": {
197
- "samples": 30,
198
- "p50_ms": 141.064813,
199
- "p95_ms": 141.96145435,
200
- "requests_per_second_from_median": 7.088940032125517,
201
- "questions_per_second_from_median": 113.42304051400828,
202
- "measured_requests_per_second": 7.0981910052955834,
203
- "measured_questions_per_second": 113.57105608472934,
204
- "max_allocated_bytes": 8963302912,
205
- "max_reserved_bytes": 9244246016,
206
- "input_tokens": 4912
207
- },
208
- "score-q32": {
209
- "samples": 30,
210
- "p50_ms": 282.3302425,
211
- "p95_ms": 285.68878665,
212
- "requests_per_second_from_median": 3.541951408198858,
213
- "questions_per_second_from_median": 113.34244506236345,
214
- "measured_requests_per_second": 3.5352975301699447,
215
- "measured_questions_per_second": 113.12952096543823,
216
- "max_allocated_bytes": 8963302912,
217
- "max_reserved_bytes": 9244246016,
218
- "input_tokens": 9824
219
- }
220
- }
221
  },
222
- "choice_q_counts": [
 
 
223
  1,
224
  2,
225
  4,
@@ -227,14 +18,96 @@
227
  16,
228
  32
229
  ],
230
- "current_p50_ms": [
231
- 32.731553500000004,
232
- 33.3187065,
233
- 41.779055,
234
- 69.641899,
235
- 141.0197655,
236
- 282.0275835
237
- ],
238
- "source_summary_sha256": "6f50e96db05a593bb24ce1c7816f52648aba80716323b9c214347eaa1295aa5c",
239
- "display_scope": "Current release only; full historical timing evidence retained in its immutable evaluation receipt."
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
240
  }
 
1
  {
2
+ "model": "Nox",
3
+ "release": "v1.3.1",
4
+ "release_binding": {
5
+ "repo_id": "llm-semantic-router/Decision-1.0-Nox",
6
+ "revision": "74979f7c7408325716dcb89b686d9cf94301923a",
7
+ "release_tag": "v1.3.1",
8
+ "bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
9
+ "publication_receipt_sha256": "7d680d3a1dc5919fc2975cbea9ea5650f39cdbf317d8e1ca6b2d03d5ae11bcdb"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
10
  },
11
+ "condition": "distinct_primary",
12
+ "type": "choice",
13
+ "Q": [
14
  1,
15
  2,
16
  4,
 
18
  16,
19
  32
20
  ],
21
+ "tokens_per_question": 499,
22
+ "samples_per_point": 30,
23
+ "fresh_process_blocks": 6,
24
+ "points": {
25
+ "1": {
26
+ "count": 30,
27
+ "p50_ms": 32.678454,
28
+ "p95_ms": 33.71943,
29
+ "mean_ms": 32.7633463,
30
+ "std_ms": 0.5253808729894747,
31
+ "minimum_ms": 31.98487,
32
+ "maximum_ms": 33.745036,
33
+ "questions_per_second_at_p50": 30.601202859841532,
34
+ "peak_allocated_bytes": 8675011072,
35
+ "peak_reserved_bytes": 9246343168,
36
+ "questions": 1,
37
+ "tokens_per_question": 499
38
+ },
39
+ "2": {
40
+ "count": 30,
41
+ "p50_ms": 36.370171,
42
+ "p95_ms": 36.64674145,
43
+ "mean_ms": 36.389989299999996,
44
+ "std_ms": 0.11128765575847639,
45
+ "minimum_ms": 36.268881,
46
+ "maximum_ms": 36.692332,
47
+ "questions_per_second_at_p50": 54.99011813829525,
48
+ "peak_allocated_bytes": 8746551808,
49
+ "peak_reserved_bytes": 9246343168,
50
+ "questions": 2,
51
+ "tokens_per_question": 499
52
+ },
53
+ "4": {
54
+ "count": 30,
55
+ "p50_ms": 60.415027,
56
+ "p95_ms": 60.9524315,
57
+ "mean_ms": 60.47764396666667,
58
+ "std_ms": 0.286383903775007,
59
+ "minimum_ms": 60.134886,
60
+ "maximum_ms": 61.481055,
61
+ "questions_per_second_at_p50": 66.20869341000211,
62
+ "peak_allocated_bytes": 8890681856,
63
+ "peak_reserved_bytes": 9246343168,
64
+ "questions": 4,
65
+ "tokens_per_question": 499
66
+ },
67
+ "8": {
68
+ "count": 30,
69
+ "p50_ms": 103.619237,
70
+ "p95_ms": 104.1805109,
71
+ "mean_ms": 103.640764,
72
+ "std_ms": 0.34493524425007943,
73
+ "minimum_ms": 102.972861,
74
+ "maximum_ms": 104.235019,
75
+ "questions_per_second_at_p50": 77.20574124667604,
76
+ "peak_allocated_bytes": 9178941952,
77
+ "peak_reserved_bytes": 9246343168,
78
+ "questions": 8,
79
+ "tokens_per_question": 499
80
+ },
81
+ "16": {
82
+ "count": 30,
83
+ "p50_ms": 206.9910775,
84
+ "p95_ms": 207.62369065000001,
85
+ "mean_ms": 206.9661473,
86
+ "std_ms": 0.49304760661506103,
87
+ "minimum_ms": 206.100008,
88
+ "maximum_ms": 207.740005,
89
+ "questions_per_second_at_p50": 77.2980178336431,
90
+ "peak_allocated_bytes": 9178946560,
91
+ "peak_reserved_bytes": 9246343168,
92
+ "questions": 16,
93
+ "tokens_per_question": 499
94
+ },
95
+ "32": {
96
+ "count": 30,
97
+ "p50_ms": 414.19944399999997,
98
+ "p95_ms": 415.3707452,
99
+ "mean_ms": 414.2741519666667,
100
+ "std_ms": 0.6025438072196239,
101
+ "minimum_ms": 413.162275,
102
+ "maximum_ms": 416.066377,
103
+ "questions_per_second_at_p50": 77.25746729877311,
104
+ "peak_allocated_bytes": 9178946560,
105
+ "peak_reserved_bytes": 9246343168,
106
+ "questions": 32,
107
+ "tokens_per_question": 499
108
+ }
109
+ },
110
+ "timing_summary_sha256": "501f25efd1813b566d77e98278229440b7a3b4a2358def1462c0171a33e9f3be",
111
+ "API_sha256": "273f6f10f22d5a68b8db34cfcbd35407fb43d8030f6d7cd188bcf118dd90a152",
112
+ "scope": "Only optimized path subsequently published, no old-release series. Fixed-length Python requests, excludes network and loading."
113
  }
release-manifest.json CHANGED
@@ -1,12 +1,12 @@
1
  {
2
  "format": "decision-public-release-v1",
3
- "status": "runtime-patch-assembled-before-staged-download-proof",
4
  "bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
5
- "readiness_sha256": "97c0c58b97685f95ad57a65b378f4e4f374ff3f7a7e3f7ef37c7960131981ffe",
6
- "model_card_sha256": "4ec3943c766b36eed886ac2713363dbcbc5b736b942c1e38f13cb1ab5f5758cd",
7
  "repo_id": "llm-semantic-router/Decision-1.0-Nox",
8
- "assembly_script_sha256": "7831ef5952b8256e585d4efa91cac4ce87738f4eb93c84b1cf1e7250def1f430",
9
- "original_bundle_manifest_preserved": false,
10
  "files_exclude_this_manifest": true,
11
  "files": [
12
  {
@@ -14,6 +14,11 @@
14
  "bytes": 6606,
15
  "sha256": "e9ed3c41423a6d15edbf82e14b80897fe5ef6cd454bcd9957d91b31460461718"
16
  },
 
 
 
 
 
17
  {
18
  "file": "Dockerfile.runtime",
19
  "bytes": 751,
@@ -21,8 +26,8 @@
21
  },
22
  {
23
  "file": "EVALUATION.md",
24
- "bytes": 9573,
25
- "sha256": "1c7e88ca75fe85502440dc4202de91786e3f68b8533e7901f398c5ee5352c4b6"
26
  },
27
  {
28
  "file": "LICENSE",
@@ -31,8 +36,8 @@
31
  },
32
  {
33
  "file": "MATERIALS.json",
34
- "bytes": 7717,
35
- "sha256": "1999846500939fcf37bd3dce8241bb3ce5b60e9c025cd486fa92faaa2d196422"
36
  },
37
  {
38
  "file": "NORMALIZATION_RUNTIME.md",
@@ -41,8 +46,8 @@
41
  },
42
  {
43
  "file": "QUESTION-SCALING.md",
44
- "bytes": 1555,
45
- "sha256": "fd5355fb6b69575a8812fa30cf1fcee2065182923206ab7dbe34a217db1bda9e"
46
  },
47
  {
48
  "file": "QWEN-LICENSE",
@@ -51,8 +56,8 @@
51
  },
52
  {
53
  "file": "README.md",
54
- "bytes": 5001,
55
- "sha256": "4ec3943c766b36eed886ac2713363dbcbc5b736b942c1e38f13cb1ab5f5758cd"
56
  },
57
  {
58
  "file": "RUNTIME-RELEASE.json",
@@ -69,6 +74,11 @@
69
  "bytes": 2621,
70
  "sha256": "fb76d9fc9147e15a678eb91dfc40d7039845a4eaebdecce041a5c9e76c3d3e0f"
71
  },
 
 
 
 
 
72
  {
73
  "file": "SERVING_OPTIMIZATION.json",
74
  "bytes": 1548,
@@ -79,6 +89,11 @@
79
  "bytes": 5637,
80
  "sha256": "d437bc0149bcb8c9891fbb33c5abc4c2336c3a89ab6cfa0981da5a7f1d19f1f4"
81
  },
 
 
 
 
 
82
  {
83
  "file": "USAGE.md",
84
  "bytes": 5435,
@@ -219,6 +234,21 @@
219
  "bytes": 2962868,
220
  "sha256": "213511289ce8df038d938ac470e803c427ed57f0f85cc397dd4d79964b866541"
221
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
222
  {
223
  "file": "assets/decision-question-scaling-600px.png",
224
  "bytes": 40825,
@@ -226,18 +256,33 @@
226
  },
227
  {
228
  "file": "assets/decision-question-scaling.pdf",
229
- "bytes": 31467,
230
- "sha256": "7e24268f9abf76a8fd431010f181dc43c5bc09478dbfe75b362435a9ca34cf8c"
231
  },
232
  {
233
  "file": "assets/decision-question-scaling.png",
234
- "bytes": 96655,
235
- "sha256": "b5f48c467a8c9b636a7bbfc34946ffe0f5c8e06210dcb40da706f62e09757910"
236
  },
237
  {
238
  "file": "assets/decision-question-scaling.svg",
239
- "bytes": 25016,
240
- "sha256": "f17f3d7dae22d2dc2ed61cb356e44130dd07b4f2544723e24d291538c8bd5e4f"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
241
  },
242
  {
243
  "file": "assets/readout.png",
@@ -314,11 +359,21 @@
314
  "bytes": 10529624,
315
  "sha256": "9cb6f639714e31bcb76b58eaf94af0b72ac3db9d091ebcda45d9b575efd489de"
316
  },
 
 
 
 
 
317
  {
318
  "file": "metrics/comparator-coverage.json",
319
  "bytes": 24628,
320
  "sha256": "fb4e6c8deee00798056aff95ceaaec4f074effbd8b84c20e69667869b1a95c03"
321
  },
 
 
 
 
 
322
  {
323
  "file": "metrics/expanded-quality.json",
324
  "bytes": 64295,
@@ -336,8 +391,8 @@
336
  },
337
  {
338
  "file": "metrics/question-scaling.json",
339
- "bytes": 9034,
340
- "sha256": "a48d1f85555426edecfcfc3a6988c5a87c4f53ef5ec10e86fd489bba50af4191"
341
  },
342
  {
343
  "file": "metrics/semantic-consistency.json",
@@ -502,9 +557,9 @@
502
  "MATERIALS.json": "copy",
503
  "metrics/semantic-consistency.json": "copy"
504
  },
505
- "scope": "Inference weights, original numerical code, calibrated runtime metadata, wrapper, model card, license, attribution and approved assets only.",
506
  "release_tag": "v1.3.1",
507
- "change_kind": "runtime_only",
508
- "previous_main_revision": "3d42ac931bba5726374d9c898e5df70db44ffea1",
509
- "previous_release_manifest_sha256": "c8a9ce88ed6312a7a822c0ef6d70f7ebceb06fadb323543fdfe494b8bdbba50d"
510
  }
 
1
  {
2
  "format": "decision-public-release-v1",
3
+ "status": "documentation-only-assembled",
4
  "bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
5
+ "readiness_sha256": "fcbb84a9bf7294ac79e2421cdd2cbfdfc46a79fd74c41ba0965778472d1cafc1",
6
+ "model_card_sha256": "108a5727cda42f93dcc2a3d5fbe653cc42a5fccbe1574ad940553aafe2f1741d",
7
  "repo_id": "llm-semantic-router/Decision-1.0-Nox",
8
+ "assembly_script_sha256": "1697c4abaee22bb7b60259c042488178fc71847afe4810037683f874d7b8d513",
9
+ "original_bundle_manifest_preserved": true,
10
  "files_exclude_this_manifest": true,
11
  "files": [
12
  {
 
14
  "bytes": 6606,
15
  "sha256": "e9ed3c41423a6d15edbf82e14b80897fe5ef6cd454bcd9957d91b31460461718"
16
  },
17
+ {
18
+ "file": "DIAGNOSTICS.md",
19
+ "bytes": 5967,
20
+ "sha256": "5492ed49869e45e45650c2501d02a78f9a2fbc17267c54f90cf034a01ebda2b3"
21
+ },
22
  {
23
  "file": "Dockerfile.runtime",
24
  "bytes": 751,
 
26
  },
27
  {
28
  "file": "EVALUATION.md",
29
+ "bytes": 3744,
30
+ "sha256": "3866dea6ae78b09bf93ad7b843d067960493217ae84b6ded87bfdac97b424608"
31
  },
32
  {
33
  "file": "LICENSE",
 
36
  },
37
  {
38
  "file": "MATERIALS.json",
39
+ "bytes": 2688,
40
+ "sha256": "387ae91720084aefff2ce1964d30550485666587e7ea3a0a2bb5fae5af678d8d"
41
  },
42
  {
43
  "file": "NORMALIZATION_RUNTIME.md",
 
46
  },
47
  {
48
  "file": "QUESTION-SCALING.md",
49
+ "bytes": 1274,
50
+ "sha256": "898f3554434284ff506dbc63c34859750f07b33dcf71f5b6d4ec7a798a56eb28"
51
  },
52
  {
53
  "file": "QWEN-LICENSE",
 
56
  },
57
  {
58
  "file": "README.md",
59
+ "bytes": 4737,
60
+ "sha256": "108a5727cda42f93dcc2a3d5fbe653cc42a5fccbe1574ad940553aafe2f1741d"
61
  },
62
  {
63
  "file": "RUNTIME-RELEASE.json",
 
74
  "bytes": 2621,
75
  "sha256": "fb76d9fc9147e15a678eb91dfc40d7039845a4eaebdecce041a5c9e76c3d3e0f"
76
  },
77
+ {
78
+ "file": "SENSITIVITY.md",
79
+ "bytes": 783,
80
+ "sha256": "e03ec18140d5c18e606da23cc5b26ef41176ab38019912d85050677e4637f0a0"
81
+ },
82
  {
83
  "file": "SERVING_OPTIMIZATION.json",
84
  "bytes": 1548,
 
89
  "bytes": 5637,
90
  "sha256": "d437bc0149bcb8c9891fbb33c5abc4c2336c3a89ab6cfa0981da5a7f1d19f1f4"
91
  },
92
+ {
93
+ "file": "TASKS.md",
94
+ "bytes": 9784,
95
+ "sha256": "e87c5271049f09ce23fcb7fdcbaa9f03e5c50039b6d7a4005a6b4c7fbc1d30ca"
96
+ },
97
  {
98
  "file": "USAGE.md",
99
  "bytes": 5435,
 
234
  "bytes": 2962868,
235
  "sha256": "213511289ce8df038d938ac470e803c427ed57f0f85cc397dd4d79964b866541"
236
  },
237
+ {
238
+ "file": "assets/decision-matrix.pdf",
239
+ "bytes": 28634,
240
+ "sha256": "d1651d91db893ffbdf3261ec476241eb03c18c7a769d19978c668d60a1ac844f"
241
+ },
242
+ {
243
+ "file": "assets/decision-matrix.png",
244
+ "bytes": 345190,
245
+ "sha256": "b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc"
246
+ },
247
+ {
248
+ "file": "assets/decision-matrix.svg",
249
+ "bytes": 46970,
250
+ "sha256": "dc92d1e3f3cf0518ee44f1bc512aae74732e1db82b919774a52c0da7c744b0b4"
251
+ },
252
  {
253
  "file": "assets/decision-question-scaling-600px.png",
254
  "bytes": 40825,
 
256
  },
257
  {
258
  "file": "assets/decision-question-scaling.pdf",
259
+ "bytes": 18958,
260
+ "sha256": "d756f60c142ecaf647e59ea5bb3a611a39d1572a6471cd49afb7ef602a23fb3b"
261
  },
262
  {
263
  "file": "assets/decision-question-scaling.png",
264
+ "bytes": 103662,
265
+ "sha256": "722e664974b28324693f5f4af6609330c278385ac118ed6f757b8fc0afec64ac"
266
  },
267
  {
268
  "file": "assets/decision-question-scaling.svg",
269
+ "bytes": 9265,
270
+ "sha256": "cc591eafa335b5d16edd306d6368679afd283a4b14d5f83fc77780a352b1d315"
271
+ },
272
+ {
273
+ "file": "assets/decision-ranking.pdf",
274
+ "bytes": 24599,
275
+ "sha256": "f3e037ada6c30871d9ebe0087fa3d122346f7ab1c927515878d524c492b7799a"
276
+ },
277
+ {
278
+ "file": "assets/decision-ranking.png",
279
+ "bytes": 226418,
280
+ "sha256": "1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0"
281
+ },
282
+ {
283
+ "file": "assets/decision-ranking.svg",
284
+ "bytes": 14935,
285
+ "sha256": "c8de982ba51fb9d2b729191018220964426a0f524a00a796553b8879db271d1a"
286
  },
287
  {
288
  "file": "assets/readout.png",
 
359
  "bytes": 10529624,
360
  "sha256": "9cb6f639714e31bcb76b58eaf94af0b72ac3db9d091ebcda45d9b575efd489de"
361
  },
362
+ {
363
+ "file": "metrics/benchmark.json",
364
+ "bytes": 2451569,
365
+ "sha256": "50dac7884849cbdb8be58c1c82b811f51543a08a548253cb8295f4e6444c4718"
366
+ },
367
  {
368
  "file": "metrics/comparator-coverage.json",
369
  "bytes": 24628,
370
  "sha256": "fb4e6c8deee00798056aff95ceaaec4f074effbd8b84c20e69667869b1a95c03"
371
  },
372
+ {
373
+ "file": "metrics/evaluation-provenance.json",
374
+ "bytes": 44403,
375
+ "sha256": "1f9c8b1e03d521160a8426798394418748daa72014dfc7272578d1e83fe237b4"
376
+ },
377
  {
378
  "file": "metrics/expanded-quality.json",
379
  "bytes": 64295,
 
391
  },
392
  {
393
  "file": "metrics/question-scaling.json",
394
+ "bytes": 3484,
395
+ "sha256": "9407f7ebf86581cdea21e704cdbebb9ced0a485eb980925f48233ebea0f830dc"
396
  },
397
  {
398
  "file": "metrics/semantic-consistency.json",
 
557
  "MATERIALS.json": "copy",
558
  "metrics/semantic-consistency.json": "copy"
559
  },
560
+ "scope": "Current14model five-panel product comparison,54task details, separate diagnostics,current API latency. Weights, runtime, tokenizer and calibration unchanged.",
561
  "release_tag": "v1.3.1",
562
+ "change_kind": "documentation-only-current-benchmark",
563
+ "previous_main_revision": "74979f7c7408325716dcb89b686d9cf94301923a",
564
+ "previous_release_manifest_sha256": "7f54061edbe5e3801107a3792d07aaf94bf51e6e1b176b3e457c86b3477134ee"
565
  }