Publish complete decision benchmark, task diagnostics, and current latency
Browse files- .gitattributes +2 -0
- DIAGNOSTICS.md +108 -0
- EVALUATION.md +26 -112
- MATERIALS.json +34 -222
- QUESTION-SCALING.md +11 -33
- README.md +27 -46
- SENSITIVITY.md +18 -0
- TASKS.md +97 -0
- assets/decision-matrix.pdf +0 -0
- assets/decision-matrix.png +3 -0
- assets/decision-matrix.svg +1253 -0
- assets/decision-question-scaling.pdf +0 -0
- assets/decision-question-scaling.png +2 -2
- assets/decision-question-scaling.svg +115 -99
- assets/decision-ranking.pdf +0 -0
- assets/decision-ranking.png +3 -0
- assets/decision-ranking.svg +352 -0
- metrics/benchmark.json +0 -0
- metrics/evaluation-provenance.json +1019 -0
- metrics/question-scaling.json +103 -230
- release-manifest.json +80 -25
.gitattributes
CHANGED
|
@@ -53,3 +53,5 @@ assets/decision-expanded-v3_core.png filter=lfs diff=lfs merge=lfs -text
|
|
| 53 |
assets/decision-expanded-v4.png filter=lfs diff=lfs merge=lfs -text
|
| 54 |
assets/decision-expanded-v5.png filter=lfs diff=lfs merge=lfs -text
|
| 55 |
assets/decision-question-scaling.png filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 53 |
assets/decision-expanded-v4.png filter=lfs diff=lfs merge=lfs -text
|
| 54 |
assets/decision-expanded-v5.png filter=lfs diff=lfs merge=lfs -text
|
| 55 |
assets/decision-question-scaling.png filter=lfs diff=lfs merge=lfs -text
|
| 56 |
+
assets/decision-matrix.png filter=lfs diff=lfs merge=lfs -text
|
| 57 |
+
assets/decision-ranking.png filter=lfs diff=lfs merge=lfs -text
|
DIAGNOSTICS.md
ADDED
|
@@ -0,0 +1,108 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Diagnostics
|
| 2 |
+
|
| 3 |
+
These axes remain separate from headline accuracy. Probability metrics use the shipped distribution; no temperature is fitted on this benchmark. An unavailable full-population metric is shown as —, without replacing it with a successful-only average.
|
| 4 |
+
|
| 5 |
+
## Probability quality on transfer
|
| 6 |
+
|
| 7 |
+
| Model | Valid / requested | Brier ↓ | NLL ↓ | ECE % ↓ | Coverage at ≤5% error % ↑ | AURC ↓ |
|
| 8 |
+
|---|---:|---:|---:|---:|---:|---:|
|
| 9 |
+
| Jev | 1046/1046 | 0.1912 | 0.5743 | 3.35 | 76.96 | 0.0346 |
|
| 10 |
+
| Lux | 1046/1046 | 0.3058 | 0.6201 | 5.13 | 57.36 | 0.0639 |
|
| 11 |
+
| Nox | 1046/1046 | 0.4350 | 0.9375 | 10.27 | 35.18 | 0.1195 |
|
| 12 |
+
| Kev-9B | 1046/1046 | 0.2948 | 0.6028 | 4.45 | 58.03 | 0.0586 |
|
| 13 |
+
| Kev-4B | 1046/1046 | 0.3145 | 0.6434 | 2.61 | 56.79 | 0.0659 |
|
| 14 |
+
| Qwen3.5-9B | 1046/1046 | 0.3673 | 0.7597 | 9.75 | 47.42 | 0.0872 |
|
| 15 |
+
| Decider | 1046/1046 | 0.4122 | 0.7969 | 8.09 | 35.18 | 0.1172 |
|
| 16 |
+
| Qwen3.5-4B | 1046/1046 | 0.4199 | 0.8186 | 8.98 | 34.70 | 0.1233 |
|
| 17 |
+
| Sol | 1046/1046 | 0.5454 | 1.1094 | 11.25 | 16.16 | 0.2073 |
|
| 18 |
+
| Kev-0.8B | 1046/1046 | 0.4839 | 0.9489 | 2.58 | 21.61 | 0.1782 |
|
| 19 |
+
| Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
|
| 20 |
+
| Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
|
| 21 |
+
| Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
|
| 22 |
+
| Kai | 1046/1046 | 0.6066 | 1.1552 | 7.95 | 1.24 | 0.3343 |
|
| 23 |
+
|
| 24 |
+
Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
|
| 25 |
+
|
| 26 |
+
## Option-order robustness
|
| 27 |
+
|
| 28 |
+
| Model | Valid / requested pairs | Both correct % ↑ | Semantic flip % ↓ | Mean half-L1 ↓ |
|
| 29 |
+
|---|---:|---:|---:|---:|
|
| 30 |
+
| Jev | 36/36 | 86.11 | 0.00 | 0.0208 |
|
| 31 |
+
| Lux | 36/36 | 77.78 | 8.33 | 0.0923 |
|
| 32 |
+
| Nox | 36/36 | 58.33 | 25.00 | 0.1049 |
|
| 33 |
+
| Kev-9B | 36/36 | 80.56 | 2.78 | 0.0615 |
|
| 34 |
+
| Kev-4B | 36/36 | 77.78 | 5.56 | 0.0667 |
|
| 35 |
+
| Qwen3.5-9B | 36/36 | 75.00 | 11.11 | 0.1157 |
|
| 36 |
+
| Decider | 36/36 | 83.33 | 11.11 | 0.0701 |
|
| 37 |
+
| Qwen3.5-4B | 36/36 | 77.78 | 13.89 | 0.1390 |
|
| 38 |
+
| Sol | 36/36 | 41.67 | 38.89 | 0.0694 |
|
| 39 |
+
| Kev-0.8B | 36/36 | 55.56 | 16.67 | 0.0808 |
|
| 40 |
+
| Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
|
| 41 |
+
| Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
|
| 42 |
+
| Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
|
| 43 |
+
| Kai | 36/36 | 41.67 | 27.78 | 0.1170 |
|
| 44 |
+
|
| 45 |
+
The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
|
| 46 |
+
|
| 47 |
+
## Missing evidence
|
| 48 |
+
|
| 49 |
+
| Model | Valid / requested | Intact/control accuracy % ↑ | Mean max P % ↓ | P≥0.9 share % ↓ | Normalized entropy ↑ | Paired confidence drop pp ↑ |
|
| 50 |
+
|---|---:|---:|---:|---:|---:|---:|
|
| 51 |
+
| Jev | 110/110 | 93.64 | 62.04 | 16.36 | 0.7657 | 30.30 |
|
| 52 |
+
| Lux | 110/110 | 84.55 | 68.81 | 19.09 | 0.6785 | 19.59 |
|
| 53 |
+
| Nox | 110/110 | 72.73 | 78.65 | 27.27 | 0.5072 | 12.64 |
|
| 54 |
+
| Kev-9B | 110/110 | 91.82 | 39.61 | 0.00 | 0.9981 | 53.19 |
|
| 55 |
+
| Kev-4B | 110/110 | 91.82 | 41.47 | 0.00 | 0.9918 | 51.57 |
|
| 56 |
+
| Qwen3.5-9B | 110/110 | 80.91 | 67.62 | 7.27 | 0.7585 | 19.09 |
|
| 57 |
+
| Decider | 110/110 | 69.09 | 68.82 | 16.36 | 0.6772 | 13.59 |
|
| 58 |
+
| Qwen3.5-4B | 110/110 | 72.73 | 59.89 | 0.91 | 0.8120 | 22.35 |
|
| 59 |
+
| Sol | 110/110 | 65.45 | 77.19 | 24.55 | 0.5424 | 5.68 |
|
| 60 |
+
| Kev-0.8B | 110/110 | 87.27 | 42.11 | 0.00 | 0.9796 | 40.34 |
|
| 61 |
+
| Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
|
| 62 |
+
| Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
|
| 63 |
+
| Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
|
| 64 |
+
| Kai | 110/110 | 37.27 | 55.39 | 18.18 | 0.8262 | -0.07 |
|
| 65 |
+
|
| 66 |
+
The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
|
| 67 |
+
|
| 68 |
+
## Native contract coverage
|
| 69 |
+
|
| 70 |
+
| Model | Original probability rows | Transfer probability rows | Transfer truncated questions |
|
| 71 |
+
|---|---:|---:|---:|
|
| 72 |
+
| Jev | 2720/2720 | 1264/1264 | — / not observable |
|
| 73 |
+
| Lux | 2720/2720 | 1264/1264 | 0 |
|
| 74 |
+
| Nox | 2720/2720 | 1264/1264 | 0 |
|
| 75 |
+
| Kev-9B | 2720/2720 | 1264/1264 | 0 |
|
| 76 |
+
| Kev-4B | 2720/2720 | 1264/1264 | 0 |
|
| 77 |
+
| Qwen3.5-9B | 2720/2720 | 1264/1264 | 0 |
|
| 78 |
+
| Decider | 2720/2720 | 1264/1264 | 0 |
|
| 79 |
+
| Qwen3.5-4B | 2720/2720 | 1264/1264 | 0 |
|
| 80 |
+
| Sol | 2720/2720 | 1264/1264 | 0 |
|
| 81 |
+
| Kev-0.8B | 2720/2720 | 1264/1264 | 0 |
|
| 82 |
+
| Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
|
| 83 |
+
| Laya · English | 2720/2720 | 1264/1264 | 34 |
|
| 84 |
+
| Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
|
| 85 |
+
| Kai | 2720/2720 | 1264/1264 | 0 |
|
| 86 |
+
|
| 87 |
+
Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation; Kai retains the shipped complete-request 1,024-token limit. Server-side truncation for Jev cannot be observed.
|
| 88 |
+
|
| 89 |
+
## Uncertainty
|
| 90 |
+
|
| 91 |
+
| Model | Overall % | 95% component-bootstrap interval |
|
| 92 |
+
|---|---:|---:|
|
| 93 |
+
| Jev | 81.05 | 79.70–82.35 |
|
| 94 |
+
| Lux | 76.72 | 75.35–78.07 |
|
| 95 |
+
| Nox | 72.84 | 71.33–74.31 |
|
| 96 |
+
| Kev-9B | 71.89 | 70.42–73.35 |
|
| 97 |
+
| Kev-4B | 70.09 | 68.45–71.63 |
|
| 98 |
+
| Qwen3.5-9B | 69.73 | 68.27–71.20 |
|
| 99 |
+
| Decider | 67.71 | 66.10–69.34 |
|
| 100 |
+
| Qwen3.5-4B | 67.29 | 65.89–68.69 |
|
| 101 |
+
| Sol | 66.32 | 64.77–67.85 |
|
| 102 |
+
| Kev-0.8B | 58.28 | 56.64–59.89 |
|
| 103 |
+
| Qwen3.5-2B | 57.24 | 55.74–58.76 |
|
| 104 |
+
| Laya · English | 51.03 | 49.43–52.68 |
|
| 105 |
+
| Laya · Multilingual | 47.19 | 45.58–48.82 |
|
| 106 |
+
| Kai | 46.49 | 45.02–47.99 |
|
| 107 |
+
|
| 108 |
+
Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
|
EVALUATION.md
CHANGED
|
@@ -1,124 +1,38 @@
|
|
| 1 |
# Evaluation
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
|
| 7 |
-
|
| 8 |
|
| 9 |
-
|
| 10 |
|
| 11 |
-
|
| 12 |
-
|---|---|---:|---:|---:|---:|---:|
|
| 13 |
-
| Nox | 4B · v1.3 | **83.00** | **51.79** | 79.06 | **86.25** | **75.03** |
|
| 14 |
-
| Sol | 2B · v1.3 | 73.75 | 46.08 | 76.56 | 84.17 | 70.14 |
|
| 15 |
-
| Kev | 9B | 76.36 | 45.54 | 86.72 | 83.75 | 73.09 |
|
| 16 |
-
| Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 71.75 |
|
| 17 |
-
| Qwen3.5 | 4B · untuned | 69.89 | 43.33 | 87.97 | 79.79 | 70.25 |
|
| 18 |
-
| Qwen3.5 | 2B · untuned | 57.12 | 39.00 | 73.75 | 72.29 | 60.54 |
|
| 19 |
-
| Laya | Upstream default | 57.01 | 37.75 | 51.25 | 63.75 | 52.44 |
|
| 20 |
-
| Jev | 1.13.0 · frontier | 79.10 | 66.38 | 94.53 | 89.79 | 82.45 |
|
| 21 |
|
| 22 |
-
|
| 23 |
-
Bold marks Sol or Nox strictly above every external open reference in that metric; Jev and the other Decision model are excluded from the threshold. Ties are not bold.
|
| 24 |
|
| 25 |
-
##
|
| 26 |
|
| 27 |
-
|
| 28 |
-
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 29 |
-
| News classification | 85.16 | 83.59 | 88.28 | 86.72 | 84.38 | 80.47 | 91.41 | 85.16 |
|
| 30 |
-
| Boolean constraints | **93.75** | 50.00 | 92.19 | 71.88 | 62.50 | 37.50 | 43.75 | 100.00 |
|
| 31 |
-
| Entity classification | 96.43 | 95.54 | 98.21 | 98.21 | 97.32 | 92.86 | 83.93 | 96.43 |
|
| 32 |
-
| Intent routing | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 85.94 | 100.00 |
|
| 33 |
-
| Evidence placement | **100.00** | **98.96** | 50.00 | 8.33 | 72.92 | 9.38 | 95.83 | 30.21 |
|
| 34 |
-
| Ordered rubric | 89.06 | 65.62 | 96.88 | 84.38 | 90.62 | 81.25 | 18.75 | 100.00 |
|
| 35 |
-
| Relation composition | 52.08 | 37.50 | 52.08 | 54.17 | 51.04 | 48.96 | 25.00 | 56.25 |
|
| 36 |
-
| Scoped evidence | **78.12** | **79.17** | 63.54 | 47.92 | 51.04 | 36.46 | 37.50 | 89.58 |
|
| 37 |
-
| State tracking | 35.42 | 27.08 | 36.46 | 29.17 | 28.12 | 25.00 | 23.96 | 33.33 |
|
| 38 |
-
| In / out of menu | **100.00** | **100.00** | 85.94 | 59.38 | 60.94 | 62.50 | 64.06 | 100.00 |
|
| 39 |
|
| 40 |
-
##
|
| 41 |
|
| 42 |
-
|
| 43 |
-
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 44 |
-
| Record identity | 53.75 | 50.00 | 46.25 | 56.25 | 50.00 | 46.25 | 47.50 | 76.25 |
|
| 45 |
-
| Capacity assignment | 46.25 | 47.50 | 50.00 | 51.25 | 50.00 | 50.00 | 63.75 | 68.75 |
|
| 46 |
-
| Constraint assignment | **41.25** | **26.25** | 25.00 | 23.75 | 23.75 | 21.25 | 22.50 | 55.00 |
|
| 47 |
-
| Intent routing · EN | 91.67 | 90.83 | 86.67 | 90.00 | 91.67 | 70.83 | 74.17 | 91.67 |
|
| 48 |
-
| Intent routing · ZH | 87.50 | 87.50 | 85.00 | 88.33 | 86.67 | 74.17 | 73.33 | 88.33 |
|
| 49 |
-
| Multiset reconciliation | 28.75 | 28.75 | 35.00 | 23.75 | 27.50 | 27.50 | 26.25 | 58.75 |
|
| 50 |
-
| Ordered service loss | **30.00** | **25.00** | 21.25 | 22.50 | 21.25 | 20.00 | 20.00 | 42.50 |
|
| 51 |
-
| Conflicting rule closure | **33.75** | 25.00 | 28.75 | 26.25 | 25.00 | 27.50 | 18.75 | 66.25 |
|
| 52 |
-
| Temporal exclusion | 42.50 | 41.25 | 32.50 | 42.50 | 31.25 | 28.75 | 12.50 | 41.25 |
|
| 53 |
-
| Transaction recovery | **62.50** | 38.75 | 45.00 | 41.25 | 26.25 | 23.75 | 18.75 | 75.00 |
|
| 54 |
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
| Task | Nox · 4B · v1.3 | Sol · 2B · v1.3 | Kev · 9B | Decider · 2B | Qwen3.5 · 4B · untuned | Qwen3.5 · 2B · untuned | Laya · Upstream default | Jev · 1.13.0 · frontier |
|
| 58 |
-
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 59 |
-
| Yes / no reading | 86.25 | 83.12 | 90.62 | 91.88 | 83.75 | 65.00 | 69.38 | 92.50 |
|
| 60 |
-
| Reading · EN | 73.12 | 70.00 | 83.75 | 93.12 | 93.12 | 83.75 | 38.12 | 96.88 |
|
| 61 |
-
| Reading · ZH | 70.62 | 70.00 | 81.88 | 91.25 | 91.25 | 81.25 | 28.12 | 96.25 |
|
| 62 |
-
|
| 63 |
-
### Reading and inference
|
| 64 |
-
|
| 65 |
-
| Task | Nox · 4B · v1.3 | Sol · 2B · v1.3 | Kev · 9B | Decider · 2B | Qwen3.5 · 4B · untuned | Qwen3.5 · 2B · untuned | Laya · Upstream default | Jev · 1.13.0 · frontier |
|
| 66 |
-
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 67 |
-
| Contextual reasoning | **70.83** | **69.17** | 67.50 | 66.67 | 60.83 | 56.67 | 30.00 | 86.67 |
|
| 68 |
-
| Answerability | **88.33** | 83.33 | 80.83 | 84.17 | 81.67 | 82.50 | 57.50 | 90.83 |
|
| 69 |
-
| Textual entailment | 89.17 | 88.33 | 90.00 | 90.83 | 80.00 | 64.17 | 72.50 | 82.50 |
|
| 70 |
-
| Scientific inference | 96.67 | 95.83 | 96.67 | 95.83 | 96.67 | 85.83 | 95.00 | 99.17 |
|
| 71 |
-
|
| 72 |
-
### Probability quality
|
| 73 |
-
|
| 74 |
-
| Model | Version | Weighted Brier ↓ | Valid probability rows |
|
| 75 |
-
|---|---|---:|---:|
|
| 76 |
-
| Nox | 4B · v1.3 | **0.3323** | 2,720 / 2,720 |
|
| 77 |
-
| Sol | 2B · v1.3 | 0.4258 | 2,720 / 2,720 |
|
| 78 |
-
| Kev | 9B | 0.3715 | 2,720 / 2,720 |
|
| 79 |
-
| Decider | 2B | 0.3535 | 2,720 / 2,720 |
|
| 80 |
-
| Qwen3.5 | 4B · untuned | 0.3996 | 2,720 / 2,720 |
|
| 81 |
-
| Qwen3.5 | 2B · untuned | 0.5130 | 2,720 / 2,720 |
|
| 82 |
-
| Laya | Upstream default | 0.5903 | 2,720 / 2,720 |
|
| 83 |
-
| Jev | 1.13.0 · frontier | 0.2287 | 2,720 / 2,720 |
|
| 84 |
-
|
| 85 |
-
Missing Brier means probability coverage was incomplete; no supported-only average is substituted.
|
| 86 |
-
The two native contract supplements are reported separately and are outside this quality mean.
|
| 87 |
-
|
| 88 |
-
## Current probability quality
|
| 89 |
-
|
| 90 |
-
Nox v1.3 raw T=1 weighted Brier is 0.349587; the shipped CAL-fitted Brier is 0.332289. Calibration uses the fixed development CAL split after weights are selected. It does not improve category accuracy by itself, and confidence is not a factuality guarantee.
|
| 91 |
-
|
| 92 |
-
## Latency
|
| 93 |
-
|
| 94 |
-
Only this release is displayed in [the latency tables and curve](QUESTION-SCALING.md). Tokenization and inference are included; loading and network are excluded. Three fresh processes supply 30 measured requests per point on an otherwise idle AMD gfx942 GPU. These measurements do not imply service concurrency or a cross-model throughput ranking.
|
| 95 |
-
|
| 96 |
-
## Comparator identities and scope
|
| 97 |
-
|
| 98 |
-
[Nox v1.3](https://huggingface.co/llm-semantic-router/Decision-1.0-Nox/tree/ad089ad3a5dc9a7a21e6d96db546bb53e2212654) and [Sol v1.3](https://huggingface.co/llm-semantic-router/Decision-1.0-Sol/tree/2412d9470d3aa125b346262dad80f6161847aa09) use their actually published bundles and completed four-panel evaluations.
|
| 99 |
-
|
| 100 |
-
[Kev-9B](https://huggingface.co/jaredpalmer/kev-9b/tree/6281032426a9ca3a08374d137bf9dfd8afd78e8a) uses its unchanged native adapter at source revision `e0bcf50153f1bda4ca6a8be5e12cbd5f9ebbce1c` under the recorded matched-FLA runtime. Its four quality panels have complete 2,720-row coverage; its separate native panels each succeed on 217 of 220 requests, so complete quality coverage must not be read as universal native-interface support.
|
| 101 |
-
|
| 102 |
-
Untuned Qwen references use the frozen chat/LM-head adapter. Laya retains its upstream default language router and native truncation; Decider retains its native adapter. Jev is the recorded 1.13.0 service snapshot; its server-side truncation is not observable. These previously observed evaluation suites are not all pristine holdouts. Supplemental finite-probe scores are outside this mean and do not alter the release rule.
|
| 103 |
-
|
| 104 |
-
The architecture, weights, tokenizer, normalization profile, temperature and original qualification are unchanged by this documentation refresh. Historical paired-update evidence remains archived as machine-readable provenance; no old release is presented as a current comparison row.
|
| 105 |
-
|
| 106 |
-
[Exact displayed statistics](metrics/expanded-quality.json) · [Comparator provenance](metrics/comparator-coverage.json) · [Figure evidence](metrics/figure-evidence.json)
|
| 107 |
-
|
| 108 |
-
## Semantic consistency
|
| 109 |
-
|
| 110 |
-
| Model | Consistent / eligible groups | Consistency ↑ |
|
| 111 |
-
|---|---:|---:|
|
| 112 |
-
| Nox · 4B · v1.3 | 174/192 | **90.62** |
|
| 113 |
-
| Sol · 2B · v1.3 | 167/192 | **86.98** |
|
| 114 |
-
| Kev · 9B | 143/192 | 74.48 |
|
| 115 |
-
| Decider · 2B | 126/192 | 65.62 |
|
| 116 |
-
| Qwen3.5 · 4B · untuned | 98/192 | 51.04 |
|
| 117 |
-
| Qwen3.5 · 2B · untuned | 94/192 | 48.96 |
|
| 118 |
-
| Laya · Upstream default | 115/192 | 59.90 |
|
| 119 |
-
| Jev · 1.13.0 · frontier | 158/192 | 82.29 |
|
| 120 |
-
|
| 121 |
-
Mixed equivalent variants of language, keys and presentation; **consistency is not accuracy**. A consistently wrong answer still counts. This observed-core supplement weights 192 eligible semantic groups equally and is outside the four-panel mean. [Definitions and exact counts](metrics/semantic-consistency.json).
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
Eligible groups require more than one variant, uniquely identifiable option semantics, and a shared semantic gold. All outputs must be valid and agree semantically. The 240 singleton news/entity groups are excluded. This is a mixed-variation measure, not an isolated option-order experiment.
|
|
|
|
| 1 |
# Evaluation
|
| 2 |
|
| 3 |
+
The comparison covers **3,766 scored decisions across 54 tasks** and all 14 displayed models. All models use the same requested examples, candidate order and scoring rules. Each model keeps its native admission, truncation and probability semantics.
|
| 4 |
|
| 5 |
+
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|
| 6 |
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 7 |
+
| Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
|
| 8 |
+
| Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
|
| 9 |
+
| Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** |
|
| 10 |
+
| Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
|
| 11 |
+
| Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
|
| 12 |
+
| Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
|
| 13 |
+
| Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 |
|
| 14 |
+
| Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 |
|
| 15 |
+
| Sol | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 |
|
| 16 |
+
| Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 |
|
| 17 |
+
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
|
| 18 |
+
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
|
| 19 |
+
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
|
| 20 |
+
| Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
|
| 21 |
|
| 22 |
+
Accuracy (%). Rows are sorted by overall score. Bold marks Decision-family cells strictly above every external open or untuned reference in that metric; Jev and the other Decision models are excluded from that threshold.
|
| 23 |
|
| 24 |
+
## Scope and weighting
|
| 25 |
|
| 26 |
+
The overall score weights **Decisions 30%, Composition 25%, Reading 15%, Inference 15%, Transfer 15%**. The first four panels contain 880, 880, 480 and 480 questions and retain their original family/source weights. Transfer is micro-accuracy over 1,046 clean knowable questions from the frozen upstream transfer test; 110 missing-evidence questions and 108 variants remain separate diagnostics. No latency, calibration error or consistency score is averaged into accuracy.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 27 |
|
| 28 |
+
These are **outcome-informed product-priority weights, chosen after observing benchmark results**. The data are observed regression tests, not a fresh blind test. Reweighting is not a training improvement. [Weight sensitivity](SENSITIVITY.md) retains the prior weighting and original four-panel comparison for the same model weights. Training, checkpoint selection and calibration do not use these test labels.
|
|
|
|
| 29 |
|
| 30 |
+
## Full results
|
| 31 |
|
| 32 |
+
[All 54 task rows](TASKS.md) preserve every original decision, composition, reading and inference task plus all 27 transfer tasks. [Diagnostics](DIAGNOSTICS.md) separately report probability quality, option-order sensitivity, missing evidence, native coverage and uncertainty. [Exact statistics](metrics/benchmark.json) include counts and confidence intervals.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
|
| 34 |
+
## Model and API scope
|
| 35 |
|
| 36 |
+
Decision models return typed Choice, Noul and Score answers in the SystemOne format. Encoder and decoder references are evaluated through their published native interfaces. Untuned Qwen models use the frozen letter-logit readout, with no generated-text parsing. Kev uses the matched BF16 backbone with FP32 head and shipped temperature; date-fact injection is disabled. This differs from the authors’ default FP32 reproduction. Laya English and multilingual are measured separately. Kai uses its shipped 1,024-token complete-request limit. Jev is a recorded hosted-service snapshot.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
|
| 38 |
+
The tests assess decisions from the supplied state and fixed choices. They do not establish live fact retrieval, universally superior reasoning or cross-hardware speed rankings. [Immutable identities and evaluation provenance](metrics/evaluation-provenance.json).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
MATERIALS.json
CHANGED
|
@@ -1,224 +1,36 @@
|
|
| 1 |
{
|
| 2 |
-
"status": "docs-only-
|
| 3 |
-
"
|
| 4 |
-
"
|
| 5 |
-
"
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
{
|
| 37 |
-
"file": "assets/decision-expanded-old_core.pdf",
|
| 38 |
-
"bytes": 44396,
|
| 39 |
-
"sha256": "59aee075c911796b81a32397011acc2de4b9825e161587dd36789f4cb99ceeab"
|
| 40 |
-
},
|
| 41 |
-
{
|
| 42 |
-
"file": "assets/decision-expanded-old_core.png",
|
| 43 |
-
"bytes": 312207,
|
| 44 |
-
"sha256": "e662de90db73e029550f5badd9a32cc2adb2e3186eb96b12940994bb3ed9eac1"
|
| 45 |
-
},
|
| 46 |
-
{
|
| 47 |
-
"file": "assets/decision-expanded-old_core.svg",
|
| 48 |
-
"bytes": 79821,
|
| 49 |
-
"sha256": "b63bef1e8d30d9ba15e04778fdff224a70d34606e552619f05c548f221cdfe04"
|
| 50 |
-
},
|
| 51 |
-
{
|
| 52 |
-
"file": "assets/decision-expanded-overview-600px.png",
|
| 53 |
-
"bytes": 94909,
|
| 54 |
-
"sha256": "f3cab743b7f82cac19857540dd773cdc25ee4db1a5f2b2671e20ff1c1a18ad19"
|
| 55 |
-
},
|
| 56 |
-
{
|
| 57 |
-
"file": "assets/decision-expanded-overview.pdf",
|
| 58 |
-
"bytes": 41559,
|
| 59 |
-
"sha256": "3734c6df527eed3988a8b14643f225c7c733a3508267305d9681ab3be26da77b"
|
| 60 |
-
},
|
| 61 |
-
{
|
| 62 |
-
"file": "assets/decision-expanded-overview.png",
|
| 63 |
-
"bytes": 187333,
|
| 64 |
-
"sha256": "c671be3a92da5bb920864d7b0cdbd6ec22e24213d66052e165c20363b319e7e5"
|
| 65 |
-
},
|
| 66 |
-
{
|
| 67 |
-
"file": "assets/decision-expanded-overview.svg",
|
| 68 |
-
"bytes": 48619,
|
| 69 |
-
"sha256": "87f8f3cf3422548d540b52a943b466497db5eab935d431c36860ad442b059ef4"
|
| 70 |
-
},
|
| 71 |
-
{
|
| 72 |
-
"file": "assets/decision-expanded-ranking-600px.png",
|
| 73 |
-
"bytes": 83252,
|
| 74 |
-
"sha256": "ce7dcccc65def38efb9822363d6be738e7ff4faaaa7d7a81242f8a42b851edc0"
|
| 75 |
-
},
|
| 76 |
-
{
|
| 77 |
-
"file": "assets/decision-expanded-ranking.pdf",
|
| 78 |
-
"bytes": 37133,
|
| 79 |
-
"sha256": "e21b19bdb952274ea9e258f5616136195bc72693709592fbf723187bf5e9dba8"
|
| 80 |
-
},
|
| 81 |
-
{
|
| 82 |
-
"file": "assets/decision-expanded-ranking.png",
|
| 83 |
-
"bytes": 186197,
|
| 84 |
-
"sha256": "e098abf598e3aa51dbef517ed7aeb6c2f7d9312d13bef37d5dbd40ce1dd76aa4"
|
| 85 |
-
},
|
| 86 |
-
{
|
| 87 |
-
"file": "assets/decision-expanded-ranking.svg",
|
| 88 |
-
"bytes": 38118,
|
| 89 |
-
"sha256": "e327283248812f9104487397ec201b28eef4f1172b058a16286ad224d884ba71"
|
| 90 |
-
},
|
| 91 |
-
{
|
| 92 |
-
"file": "assets/decision-expanded-v3_core-600px.png",
|
| 93 |
-
"bytes": 168294,
|
| 94 |
-
"sha256": "0630ac755fbfb100ebb3ecb965a5ebe34f511384b1a4c98468d6f0fdbc239d13"
|
| 95 |
-
},
|
| 96 |
-
{
|
| 97 |
-
"file": "assets/decision-expanded-v3_core.pdf",
|
| 98 |
-
"bytes": 44466,
|
| 99 |
-
"sha256": "426b16796dd3efaf4dc9e50ffd081ec0707831e391629b0342734b3ee334d28e"
|
| 100 |
-
},
|
| 101 |
-
{
|
| 102 |
-
"file": "assets/decision-expanded-v3_core.png",
|
| 103 |
-
"bytes": 300404,
|
| 104 |
-
"sha256": "d5eaa30f40e03e06414650c023c4994c13182a6847ec423155a9e126df1d16a0"
|
| 105 |
-
},
|
| 106 |
-
{
|
| 107 |
-
"file": "assets/decision-expanded-v3_core.svg",
|
| 108 |
-
"bytes": 79901,
|
| 109 |
-
"sha256": "5797fa75fd89e31ec23bd47096258e976f916ab5b1adf2f6cdf19f25088f97a1"
|
| 110 |
-
},
|
| 111 |
-
{
|
| 112 |
-
"file": "assets/decision-expanded-v4-600px.png",
|
| 113 |
-
"bytes": 76655,
|
| 114 |
-
"sha256": "bbd20f118917a754a4844eae746c19e04e07a9629b4ecc8a1dc5e9e83b28c377"
|
| 115 |
-
},
|
| 116 |
-
{
|
| 117 |
-
"file": "assets/decision-expanded-v4.pdf",
|
| 118 |
-
"bytes": 37682,
|
| 119 |
-
"sha256": "e2015bc6c8b661c58fe52a3f575dab293c77050ab68145032ccaca537ae18ee8"
|
| 120 |
-
},
|
| 121 |
-
{
|
| 122 |
-
"file": "assets/decision-expanded-v4.png",
|
| 123 |
-
"bytes": 152599,
|
| 124 |
-
"sha256": "47f5d01cea0c8d620f9076f8972bf3457c7e234c99ddcb734250ea3b118e78a9"
|
| 125 |
-
},
|
| 126 |
-
{
|
| 127 |
-
"file": "assets/decision-expanded-v4.svg",
|
| 128 |
-
"bytes": 42942,
|
| 129 |
-
"sha256": "2ba77ba12ca9eb71105b2e09967a46c223235d2aa82c3175cc0982c7db8cb5b3"
|
| 130 |
-
},
|
| 131 |
-
{
|
| 132 |
-
"file": "assets/decision-expanded-v5-600px.png",
|
| 133 |
-
"bytes": 95077,
|
| 134 |
-
"sha256": "8eebb01aa741a9cc505e3e3f1841e68d443208efe98b7f9297d7514d1993cdba"
|
| 135 |
-
},
|
| 136 |
-
{
|
| 137 |
-
"file": "assets/decision-expanded-v5.pdf",
|
| 138 |
-
"bytes": 40182,
|
| 139 |
-
"sha256": "33849a46b149838e4449bc4db80926d6b13b31539884056695ae3235bb5c3190"
|
| 140 |
-
},
|
| 141 |
-
{
|
| 142 |
-
"file": "assets/decision-expanded-v5.png",
|
| 143 |
-
"bytes": 187696,
|
| 144 |
-
"sha256": "057fe35e93ebe41f3bb8653392c02de27616aa248fb6e2a25026f3d2b1ba9f5a"
|
| 145 |
-
},
|
| 146 |
-
{
|
| 147 |
-
"file": "assets/decision-expanded-v5.svg",
|
| 148 |
-
"bytes": 48613,
|
| 149 |
-
"sha256": "42f77533052f02089f27f557dee4a6990c835e36da091f4e5458198e496232b6"
|
| 150 |
-
},
|
| 151 |
-
{
|
| 152 |
-
"file": "assets/decision-family-header.png",
|
| 153 |
-
"bytes": 2962868,
|
| 154 |
-
"sha256": "213511289ce8df038d938ac470e803c427ed57f0f85cc397dd4d79964b866541"
|
| 155 |
-
},
|
| 156 |
-
{
|
| 157 |
-
"file": "assets/decision-question-scaling-600px.png",
|
| 158 |
-
"bytes": 40825,
|
| 159 |
-
"sha256": "6280e56c1af9f318def3fc13610555cfdc9b9a4af2c69aec8b4226b3cfa41963"
|
| 160 |
-
},
|
| 161 |
-
{
|
| 162 |
-
"file": "assets/decision-question-scaling.pdf",
|
| 163 |
-
"bytes": 31467,
|
| 164 |
-
"sha256": "7e24268f9abf76a8fd431010f181dc43c5bc09478dbfe75b362435a9ca34cf8c"
|
| 165 |
-
},
|
| 166 |
-
{
|
| 167 |
-
"file": "assets/decision-question-scaling.png",
|
| 168 |
-
"bytes": 96655,
|
| 169 |
-
"sha256": "b5f48c467a8c9b636a7bbfc34946ffe0f5c8e06210dcb40da706f62e09757910"
|
| 170 |
-
},
|
| 171 |
-
{
|
| 172 |
-
"file": "assets/decision-question-scaling.svg",
|
| 173 |
-
"bytes": 25016,
|
| 174 |
-
"sha256": "f17f3d7dae22d2dc2ed61cb356e44130dd07b4f2544723e24d291538c8bd5e4f"
|
| 175 |
-
},
|
| 176 |
-
{
|
| 177 |
-
"file": "assets/readout.png",
|
| 178 |
-
"bytes": 285481,
|
| 179 |
-
"sha256": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085"
|
| 180 |
-
},
|
| 181 |
-
{
|
| 182 |
-
"file": "metrics/comparator-coverage.json",
|
| 183 |
-
"bytes": 24628,
|
| 184 |
-
"sha256": "fb4e6c8deee00798056aff95ceaaec4f074effbd8b84c20e69667869b1a95c03"
|
| 185 |
-
},
|
| 186 |
-
{
|
| 187 |
-
"file": "metrics/expanded-quality.json",
|
| 188 |
-
"bytes": 64295,
|
| 189 |
-
"sha256": "7b8936b43bce9d19d9010ab61225a6460160b3aa76c37f5d6df5ba650b433c6c"
|
| 190 |
-
},
|
| 191 |
-
{
|
| 192 |
-
"file": "metrics/figure-evidence.json",
|
| 193 |
-
"bytes": 8746,
|
| 194 |
-
"sha256": "d9b425316f0d4c1caebb2587cb1027597d421853a089a94e330d016901c5230c"
|
| 195 |
-
},
|
| 196 |
-
{
|
| 197 |
-
"file": "metrics/materials-provenance.json",
|
| 198 |
-
"bytes": 1333,
|
| 199 |
-
"sha256": "3f052b894ab0780825ef698556483a5900d1e024cfcfaa1e89919261d812adb4"
|
| 200 |
-
},
|
| 201 |
-
{
|
| 202 |
-
"file": "metrics/question-scaling.json",
|
| 203 |
-
"bytes": 9034,
|
| 204 |
-
"sha256": "a48d1f85555426edecfcfc3a6988c5a87c4f53ef5ec10e86fd489bba50af4191"
|
| 205 |
-
},
|
| 206 |
-
{
|
| 207 |
-
"file": "metrics/semantic-consistency.json",
|
| 208 |
-
"bytes": 20784,
|
| 209 |
-
"sha256": "5a389236a22e1068be862b2daf88a5caf5a384a9260f0cafbaca9a37b18d8d26"
|
| 210 |
-
},
|
| 211 |
-
{
|
| 212 |
-
"file": "metrics/uncalibrated-quality.json",
|
| 213 |
-
"bytes": 19925,
|
| 214 |
-
"sha256": "2bc4aca25ef6468bbf3d872e0d72d50476c921717b66ad8b2469f70b051ae735"
|
| 215 |
-
},
|
| 216 |
-
{
|
| 217 |
-
"file": "model-card-example.json",
|
| 218 |
-
"bytes": 3775,
|
| 219 |
-
"sha256": "9d17a7495a0986d4850a966e282cd3d9e0a31130bdf08c0a9e60461795ce5502"
|
| 220 |
-
}
|
| 221 |
-
],
|
| 222 |
-
"post_render_finalizer_sha256": "4818899220640baf74e6f87915dfe4ca5f77aedeca40e63c6f95ff9db1cf3d3e",
|
| 223 |
-
"linear_latency_finalizer_sha256": "750bfff4297d24237896ef8080758045bc2f12c5fb7ce1a6a790fe871371c161"
|
| 224 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"status": "reviewable-docs-only-payload",
|
| 3 |
+
"family": "Nox",
|
| 4 |
+
"gate_sha256": "687163d21bc76a3e096ade8860b98e55711692e7890c9f5fe263caba1352d3e9",
|
| 5 |
+
"statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
|
| 6 |
+
"source_sha256": "92b0b306bfcb8f04c429687b6154695f6b058d168fed30eeae7c96c860bc991f",
|
| 7 |
+
"quality_method": "outcome-informed product weights; unchanged model predictions",
|
| 8 |
+
"all14_models": true,
|
| 9 |
+
"all54_tasks": true,
|
| 10 |
+
"latency_current_API_six_loads_30_samples": true,
|
| 11 |
+
"model_or_runtime_changes": false,
|
| 12 |
+
"public_secret_scan": "PASS",
|
| 13 |
+
"files": {
|
| 14 |
+
"DIAGNOSTICS.md": "5492ed49869e45e45650c2501d02a78f9a2fbc17267c54f90cf034a01ebda2b3",
|
| 15 |
+
"EVALUATION.md": "3866dea6ae78b09bf93ad7b843d067960493217ae84b6ded87bfdac97b424608",
|
| 16 |
+
"QUESTION-SCALING.md": "898f3554434284ff506dbc63c34859750f07b33dcf71f5b6d4ec7a798a56eb28",
|
| 17 |
+
"README.md": "108a5727cda42f93dcc2a3d5fbe653cc42a5fccbe1574ad940553aafe2f1741d",
|
| 18 |
+
"SENSITIVITY.md": "e03ec18140d5c18e606da23cc5b26ef41176ab38019912d85050677e4637f0a0",
|
| 19 |
+
"TASKS.md": "e87c5271049f09ce23fcb7fdcbaa9f03e5c50039b6d7a4005a6b4c7fbc1d30ca",
|
| 20 |
+
"assets/architecture.png": "8fcfa10b3941ee4aec1f5fba20a25e382a7d5de83074071c9ef68e627996de40",
|
| 21 |
+
"assets/architecture.svg": "215d9af6f94cc24847fc1d23a0b2a287b36afdc1e10ea4f5c82895b2fe103002",
|
| 22 |
+
"assets/decision-matrix.pdf": "d1651d91db893ffbdf3261ec476241eb03c18c7a769d19978c668d60a1ac844f",
|
| 23 |
+
"assets/decision-matrix.png": "b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc",
|
| 24 |
+
"assets/decision-matrix.svg": "dc92d1e3f3cf0518ee44f1bc512aae74732e1db82b919774a52c0da7c744b0b4",
|
| 25 |
+
"assets/decision-question-scaling.pdf": "d756f60c142ecaf647e59ea5bb3a611a39d1572a6471cd49afb7ef602a23fb3b",
|
| 26 |
+
"assets/decision-question-scaling.png": "722e664974b28324693f5f4af6609330c278385ac118ed6f757b8fc0afec64ac",
|
| 27 |
+
"assets/decision-question-scaling.svg": "cc591eafa335b5d16edd306d6368679afd283a4b14d5f83fc77780a352b1d315",
|
| 28 |
+
"assets/decision-ranking.pdf": "f3e037ada6c30871d9ebe0087fa3d122346f7ab1c927515878d524c492b7799a",
|
| 29 |
+
"assets/decision-ranking.png": "1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0",
|
| 30 |
+
"assets/decision-ranking.svg": "c8de982ba51fb9d2b729191018220964426a0f524a00a796553b8879db271d1a",
|
| 31 |
+
"assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
|
| 32 |
+
"metrics/benchmark.json": "50dac7884849cbdb8be58c1c82b811f51543a08a548253cb8295f4e6444c4718",
|
| 33 |
+
"metrics/evaluation-provenance.json": "1f9c8b1e03d521160a8426798394418748daa72014dfc7272578d1e83fe237b4",
|
| 34 |
+
"metrics/question-scaling.json": "9407f7ebf86581cdea21e704cdbebb9ced0a485eb980925f48233ebea0f830dc"
|
| 35 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
}
|
QUESTION-SCALING.md
CHANGED
|
@@ -1,42 +1,20 @@
|
|
| 1 |
# Request latency
|
| 2 |
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
Three independently loaded blocks provide 30 measured requests per point after three warmups per block. End-to-end Python latency includes tokenization and inference, excluding model loading and network. No cross-request prefix cache is enabled. These are sequential requests, not concurrent-service throughput.
|
| 6 |
|
| 7 |

|
| 8 |
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
| Questions | p50 ms ↓ | p95 ms ↓ | Input tokens |
|
| 12 |
-
|---:|---:|---:|---:|
|
| 13 |
-
| 1 | 32.732 | 33.474 | 309 |
|
| 14 |
-
| 2 | 33.319 | 33.560 | 618 |
|
| 15 |
-
| 4 | 41.779 | 42.215 | 1236 |
|
| 16 |
-
| 8 | 69.642 | 70.728 | 2472 |
|
| 17 |
-
| 16 | 141.020 | 141.853 | 4944 |
|
| 18 |
-
| 32 | 282.028 | 283.684 | 9888 |
|
| 19 |
-
|
| 20 |
-
## Noul
|
| 21 |
-
|
| 22 |
-
| Questions | p50 ms ↓ | p95 ms ↓ | Input tokens |
|
| 23 |
|---:|---:|---:|---:|
|
| 24 |
-
| 1 | 32.
|
| 25 |
-
| 2 |
|
| 26 |
-
| 4 |
|
| 27 |
-
| 8 |
|
| 28 |
-
| 16 |
|
| 29 |
-
| 32 |
|
| 30 |
|
| 31 |
-
|
| 32 |
|
| 33 |
-
|
| 34 |
-
|---:|---:|---:|---:|
|
| 35 |
-
| 1 | 33.319 | 34.269 | 307 |
|
| 36 |
-
| 2 | 33.356 | 33.767 | 614 |
|
| 37 |
-
| 4 | 41.859 | 42.210 | 1228 |
|
| 38 |
-
| 8 | 69.734 | 70.124 | 2456 |
|
| 39 |
-
| 16 | 141.065 | 141.961 | 4912 |
|
| 40 |
-
| 32 | 282.330 | 285.689 | 9824 |
|
| 41 |
|
| 42 |
-
|
|
|
|
| 1 |
# Request latency
|
| 2 |
|
| 3 |
+
Nox · distinct Choice questions with **499 input tokens per question**. Only the number of questions changes.
|
|
|
|
|
|
|
| 4 |
|
| 5 |

|
| 6 |
|
| 7 |
+
| Questions | p50 ms ↓ | p95 ms ↓ | Peak allocated GiB |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
|---:|---:|---:|---:|
|
| 9 |
+
| 1 | 32.678 | 33.719 | 8.079 |
|
| 10 |
+
| 2 | 36.370 | 36.647 | 8.146 |
|
| 11 |
+
| 4 | 60.415 | 60.952 | 8.280 |
|
| 12 |
+
| 8 | 103.619 | 104.181 | 8.549 |
|
| 13 |
+
| 16 | 206.991 | 207.624 | 8.549 |
|
| 14 |
+
| 32 | 414.199 | 415.371 | 8.549 |
|
| 15 |
|
| 16 |
+
Six independently loaded process blocks supply 30 measured requests per point after warmup. End-to-end Python request latency includes rendering, tokenization, inference, output assembly and final synchronization. Model loading and network are excluded. Requests are sequential, with no cross-request prefix cache.
|
| 17 |
|
| 18 |
+
The measured implementation is bound byte-for-byte to the current published API, verified by an offline Hub-download proof and full 3,160-answer regression. This documentation update changes no inference code or weights. [Exact measurements and immutable runtime identity](metrics/question-scaling.json).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
|
| 20 |
+
These fixed short-input Choice measurements do not establish concurrent HTTP throughput, long-context scaling, other question-type performance or a cross-hardware speed ranking.
|
README.md
CHANGED
|
@@ -14,13 +14,11 @@ tags:
|
|
| 14 |
- rocm
|
| 15 |
---
|
| 16 |
|
| 17 |
-

|
| 18 |
-
|
| 19 |
# Decision-1.0-Nox
|
| 20 |
|
| 21 |
*Nox, Latin for night.*
|
| 22 |
|
| 23 |
-
**Your move.** Give Nox a state, questions and possible answers. It returns decisions and probabilities with labels defined at runtime.
|
| 24 |
|
| 25 |
**4.208B parameters · 16K complete-question budget · English / Chinese evaluated · Apache 2.0**
|
| 26 |
|
|
@@ -34,59 +32,42 @@ tags:
|
|
| 34 |
|
| 35 |
## Measured capability
|
| 36 |
|
| 37 |
-
**
|
| 38 |
-
|
| 39 |
-
| Model | Mean accuracy ↑ | Decisions | Composition | Reading | Inference |
|
| 40 |
-
|---|---:|---:|---:|---:|---:|
|
| 41 |
-
| Nox · 4B · v1.3 | **75.03** | **83.00** | **51.79** | 79.06 | **86.25** |
|
| 42 |
-
| Sol · 2B · v1.3 | 70.14 | 73.75 | 46.08 | 76.56 | 84.17 |
|
| 43 |
-
| Kev · 9B | 73.09 | 76.36 | 45.54 | 86.72 | 83.75 |
|
| 44 |
-
| Decider · 2B | 71.75 | 64.01 | 46.58 | 92.03 | 84.38 |
|
| 45 |
-
| Qwen3.5 · 4B · untuned | 70.25 | 69.89 | 43.33 | 87.97 | 79.79 |
|
| 46 |
-
| Qwen3.5 · 2B · untuned | 60.54 | 57.12 | 39.00 | 73.75 | 72.29 |
|
| 47 |
-
| Laya · Upstream default | 52.44 | 57.01 | 37.75 | 51.25 | 63.75 |
|
| 48 |
-
| Jev · 1.13.0 · frontier | 82.45 | 79.10 | 66.38 | 94.53 | 89.79 |
|
| 49 |
-
|
| 50 |
-
Accuracy (%). Each panel contributes one quarter with its frozen family/source weights. **Bold** marks Sol or Nox strictly above every external open reference for that metric; the other Decision model and Jev are excluded from the threshold. Jev is a closed-service frontier reference. [Methods and uncertainty](EVALUATION.md).
|
| 51 |
-
|
| 52 |
-

|
| 53 |
-
|
| 54 |
-

|
| 55 |
-
|
| 56 |
-
## Detailed capabilities
|
| 57 |
-
|
| 58 |
-

|
| 59 |
-
|
| 60 |
-

|
| 61 |
|
| 62 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
|
| 64 |
-
|
| 65 |
|
| 66 |
-
|
| 67 |
|
| 68 |
-
|
| 69 |
-
|---|---:|---:|
|
| 70 |
-
| Nox · 4B · v1.3 | 174/192 | **90.62** |
|
| 71 |
-
| Sol · 2B · v1.3 | 167/192 | **86.98** |
|
| 72 |
-
| Kev · 9B | 143/192 | 74.48 |
|
| 73 |
-
| Decider · 2B | 126/192 | 65.62 |
|
| 74 |
-
| Qwen3.5 · 4B · untuned | 98/192 | 51.04 |
|
| 75 |
-
| Qwen3.5 · 2B · untuned | 94/192 | 48.96 |
|
| 76 |
-
| Laya · Upstream default | 115/192 | 59.90 |
|
| 77 |
-
| Jev · 1.13.0 · frontier | 158/192 | 82.29 |
|
| 78 |
|
| 79 |
-
|
| 80 |
|
| 81 |
-
## More questions,
|
| 82 |
|
| 83 |
-
![
|
| 84 |
|
| 85 |
-
|
| 86 |
|
| 87 |
## Try it
|
| 88 |
|
| 89 |
-
Download `hf download llm-semantic-router/Decision-1.0-Nox --
|
| 90 |
|
| 91 |
```python
|
| 92 |
from decision import DecisionModel
|
|
@@ -96,7 +77,7 @@ model = DecisionModel.from_pretrained("/model", local_files_only=True)
|
|
| 96 |
print(model.decide(**REQUEST)["answers"])
|
| 97 |
```
|
| 98 |
|
| 99 |
-
[Tested request and output](model-card-example.json) · [Install and API guide](USAGE.md)
|
| 100 |
|
| 101 |
The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.
|
| 102 |
|
|
|
|
| 14 |
- rocm
|
| 15 |
---
|
| 16 |
|
|
|
|
|
|
|
| 17 |
# Decision-1.0-Nox
|
| 18 |
|
| 19 |
*Nox, Latin for night.*
|
| 20 |
|
| 21 |
+
**Your move.** Give Nox a state, questions and possible answers. It returns typed decisions and probabilities, with labels defined at runtime.
|
| 22 |
|
| 23 |
**4.208B parameters · 16K complete-question budget · English / Chinese evaluated · Apache 2.0**
|
| 24 |
|
|
|
|
| 32 |
|
| 33 |
## Measured capability
|
| 34 |
|
| 35 |
+
**72.84% overall accuracy** across 3,766 scored decisions and 54 tasks. Nox leads Kev-4B by **2.75 percentage points** on this decision-focused comparison; reading and transfer remain opportunities to improve.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
|
| 37 |
+
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|
| 38 |
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 39 |
+
| Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
|
| 40 |
+
| Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
|
| 41 |
+
| Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** |
|
| 42 |
+
| Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
|
| 43 |
+
| Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
|
| 44 |
+
| Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
|
| 45 |
+
| Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 |
|
| 46 |
+
| Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 |
|
| 47 |
+
| Sol | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 |
|
| 48 |
+
| Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 |
|
| 49 |
+
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
|
| 50 |
+
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
|
| 51 |
+
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
|
| 52 |
+
| Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
|
| 53 |
|
| 54 |
+
Accuracy (%). Overall weights: Decisions **30%**, Composition **25%**, Reading **15%**, Inference **15%**, Transfer **15%**. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.
|
| 55 |
|
| 56 |
+

|
| 57 |
|
| 58 |
+

|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 59 |
|
| 60 |
+
[All 54 tasks](TASKS.md) · [Order, missing-evidence and calibration diagnostics](DIAGNOSTICS.md) · [Methods and uncertainty](EVALUATION.md)
|
| 61 |
|
| 62 |
+
## More questions, one request
|
| 63 |
|
| 64 |
+

|
| 65 |
|
| 66 |
+
Distinct Choice questions at a fixed **499 input tokens per question**. Thirty measurements per point across six independently loaded processes on an otherwise idle AMD gfx942 GPU. Python latency includes tokenization and inference; loading and network are excluded. [p50, p95 and memory](QUESTION-SCALING.md).
|
| 67 |
|
| 68 |
## Try it
|
| 69 |
|
| 70 |
+
Download the current model with `hf download llm-semantic-router/Decision-1.0-Nox --local-dir decision-model`, then follow [ROCm setup](RUNTIME.md). In that container, with the model mounted at `/model`:
|
| 71 |
|
| 72 |
```python
|
| 73 |
from decision import DecisionModel
|
|
|
|
| 77 |
print(model.decide(**REQUEST)["answers"])
|
| 78 |
```
|
| 79 |
|
| 80 |
+
[Tested SystemOne request and output](model-card-example.json) · [Install and typed API guide](USAGE.md)
|
| 81 |
|
| 82 |
The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.
|
| 83 |
|
SENSITIVITY.md
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Product weights were changed after earlier results were observed. These are identical model predictions under different weights, not trained-model improvements.
|
| 2 |
+
|
| 3 |
+
| Model | Current 30/25/15/15/15 | Earlier 25/25/15/15/20 | Original four-panel mean |
|
| 4 |
+
|---|---:|---:|---:|
|
| 5 |
+
| Jev | 81.05 | 81.45 | 82.45 |
|
| 6 |
+
| Lux | 76.72 | 76.43 | 79.06 |
|
| 7 |
+
| Nox | 72.84 | 72.09 | 75.03 |
|
| 8 |
+
| Kev-9B | 71.89 | 72.01 | 73.19 |
|
| 9 |
+
| Kev-4B | 70.09 | 70.30 | 71.73 |
|
| 10 |
+
| Qwen3.5-9B | 69.73 | 69.70 | 71.99 |
|
| 11 |
+
| Decider | 67.71 | 67.97 | 71.75 |
|
| 12 |
+
| Qwen3.5-4B | 67.29 | 67.24 | 70.25 |
|
| 13 |
+
| Sol | 66.32 | 65.48 | 70.14 |
|
| 14 |
+
| Kev-0.8B | 58.28 | 58.33 | 59.75 |
|
| 15 |
+
| Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
|
| 16 |
+
| Laya · English | 51.03 | 50.85 | 51.76 |
|
| 17 |
+
| Laya · Multilingual | 47.19 | 47.18 | 48.56 |
|
| 18 |
+
| Kai | 46.49 | 46.80 | 48.16 |
|
TASKS.md
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# All 54 tasks
|
| 2 |
+
|
| 3 |
+
Accuracy (%) on the same requested rows. Bold marks a Decision-family result strictly above every external open or untuned reference; Jev and the other Decision models do not set that threshold. Ties are not bold. Each task uses its full requested denominator. These task rows are diagnostic; the headline retains its declared within-panel weights.
|
| 4 |
+
|
| 5 |
+
<details>
|
| 6 |
+
<summary>Decisions · 10 tasks</summary>
|
| 7 |
+
|
| 8 |
+
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
|
| 9 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 10 |
+
| News classification | 128 | 85.16 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 | 78.12 |
|
| 11 |
+
| Boolean constraints | 64 | 100.00 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 | 53.12 |
|
| 12 |
+
| Entity classification | 112 | 96.43 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 | 50.00 |
|
| 13 |
+
| Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 | 81.25 |
|
| 14 |
+
| Evidence placement | 96 | 30.21 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 | 1.04 |
|
| 15 |
+
| Ordered rubric | 64 | 100.00 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 | 6.25 |
|
| 16 |
+
| Relation composition | 96 | 56.25 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 | 51.04 |
|
| 17 |
+
| Scoped evidence | 96 | 89.58 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 | 29.17 |
|
| 18 |
+
| State tracking | 96 | 33.33 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 | 30.21 |
|
| 19 |
+
| In / out of menu | 64 | 100.00 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 | 42.19 |
|
| 20 |
+
|
| 21 |
+
</details>
|
| 22 |
+
|
| 23 |
+
<details>
|
| 24 |
+
<summary>Composition · 10 tasks</summary>
|
| 25 |
+
|
| 26 |
+
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
|
| 27 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 28 |
+
| Record identity | 80 | 76.25 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 | 51.25 |
|
| 29 |
+
| Capacity assignment | 80 | 68.75 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 | 50.00 |
|
| 30 |
+
| Constraint assignment | 80 | 55.00 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 | 31.25 |
|
| 31 |
+
| Intent routing · EN | 120 | 91.67 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 | 72.50 |
|
| 32 |
+
| Intent routing · ZH | 120 | 88.33 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 | 67.50 |
|
| 33 |
+
| Multiset reconciliation | 80 | 58.75 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 | 32.50 |
|
| 34 |
+
| Ordered service loss | 80 | 42.50 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 | 21.25 |
|
| 35 |
+
| Conflicting rule closure | 80 | 66.25 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 | 23.75 |
|
| 36 |
+
| Temporal exclusion | 80 | 41.25 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 | 22.50 |
|
| 37 |
+
| Transaction recovery | 80 | 75.00 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 | 26.25 |
|
| 38 |
+
|
| 39 |
+
</details>
|
| 40 |
+
|
| 41 |
+
<details>
|
| 42 |
+
<summary>Reading · 3 tasks</summary>
|
| 43 |
+
|
| 44 |
+
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
|
| 45 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 46 |
+
| Yes / no reading | 160 | 92.50 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 | 74.38 |
|
| 47 |
+
| Reading · EN | 160 | 96.88 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 | 39.38 |
|
| 48 |
+
| Reading · ZH | 160 | 96.25 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 | 35.62 |
|
| 49 |
+
|
| 50 |
+
</details>
|
| 51 |
+
|
| 52 |
+
<details>
|
| 53 |
+
<summary>Inference · 4 tasks</summary>
|
| 54 |
+
|
| 55 |
+
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
|
| 56 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 57 |
+
| Contextual reasoning | 120 | 86.67 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 | 32.50 |
|
| 58 |
+
| Answerability | 120 | 90.83 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 | 55.83 |
|
| 59 |
+
| Textual entailment | 120 | 82.50 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 | 34.17 |
|
| 60 |
+
| Scientific inference | 120 | 99.17 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 | 95.83 |
|
| 61 |
+
|
| 62 |
+
</details>
|
| 63 |
+
|
| 64 |
+
<details>
|
| 65 |
+
<summary>Transfer · 27 tasks</summary>
|
| 66 |
+
|
| 67 |
+
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
|
| 68 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 69 |
+
| Buried emotion | 20 | 70.00 | 40.00 | 20.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 | 55.00 |
|
| 70 |
+
| Buried paraphrase | 20 | 95.00 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 | 35.00 |
|
| 71 |
+
| Buried entailment | 20 | 95.00 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 | 55.00 |
|
| 72 |
+
| Buried offensive-language detection | 20 | 55.00 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 | 75.00 |
|
| 73 |
+
| Combined policy conditions | 32 | 96.88 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 | 34.38 |
|
| 74 |
+
| Policy exceptions | 32 | 100.00 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 | 62.50 |
|
| 75 |
+
| Policy negation | 32 | 90.62 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 | 56.25 |
|
| 76 |
+
| Authorization contrast | 40 | 100.00 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 | 50.00 |
|
| 77 |
+
| Deadline contrast | 40 | 92.50 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 | 47.50 |
|
| 78 |
+
| Emotion | 80 | 67.50 | 61.25 | 40.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 | 52.50 |
|
| 79 |
+
| MMLU | 80 | 88.75 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 | 31.25 |
|
| 80 |
+
| MMLU-Pro | 200 | 84.00 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 | 17.50 |
|
| 81 |
+
| Paraphrase | 80 | 87.50 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 | 47.50 |
|
| 82 |
+
| Question entailment | 80 | 91.25 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 | 71.25 |
|
| 83 |
+
| Science questions | 80 | 100.00 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 | 92.50 |
|
| 84 |
+
| Offensive-language detection | 80 | 76.25 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 | 78.75 |
|
| 85 |
+
| Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 | 50.00 |
|
| 86 |
+
| Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 | 50.00 |
|
| 87 |
+
| Evidence control · deadline | 10 | 100.00 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 | 50.00 |
|
| 88 |
+
| Evidence control · late fee | 10 | 100.00 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 | 30.00 |
|
| 89 |
+
| Evidence control · quantity limit | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 | 50.00 |
|
| 90 |
+
| Evidence control · return window | 10 | 50.00 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 | 50.00 |
|
| 91 |
+
| Evidence control · shipping delay | 10 | 90.00 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 | 20.00 |
|
| 92 |
+
| Evidence control · sla response | 10 | 100.00 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 | 20.00 |
|
| 93 |
+
| Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 | 50.00 |
|
| 94 |
+
| Evidence control · volume discount | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 | 30.00 |
|
| 95 |
+
| Evidence control · warranty claim | 10 | 90.00 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 | 10.00 |
|
| 96 |
+
|
| 97 |
+
</details>
|
assets/decision-matrix.pdf
ADDED
|
Binary file (28.6 kB). View file
|
|
|
assets/decision-matrix.png
ADDED
|
Git LFS Details
|
assets/decision-matrix.svg
ADDED
|
|
assets/decision-question-scaling.pdf
CHANGED
|
Binary files a/assets/decision-question-scaling.pdf and b/assets/decision-question-scaling.pdf differ
|
|
|
assets/decision-question-scaling.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
assets/decision-question-scaling.svg
CHANGED
|
|
|
|
assets/decision-ranking.pdf
ADDED
|
Binary file (24.6 kB). View file
|
|
|
assets/decision-ranking.png
ADDED
|
Git LFS Details
|
assets/decision-ranking.svg
ADDED
|
|
metrics/benchmark.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
metrics/evaluation-provenance.json
ADDED
|
@@ -0,0 +1,1019 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"protocol": {
|
| 3 |
+
"version": "decision-priority-benchmark-v4",
|
| 4 |
+
"utc": "2026-09-22T05:05:12.118255+00:00",
|
| 5 |
+
"status": "User-requested product weighting, observed regression benchmark; release remains evidence-gated",
|
| 6 |
+
"authorization": "Latest user explicitly requested transfer20%\u219215% and assigned released5% to the stronger original-decision/composition area. Original decisions chosen from already observed same-size gaps. This is outcome-informed product weighting, not neutral/blind prospective benchmark design.",
|
| 7 |
+
"weights": {
|
| 8 |
+
"old_core": "3/10",
|
| 9 |
+
"v3_core": "1/4",
|
| 10 |
+
"v4": "3/20",
|
| 11 |
+
"v5": "3/20",
|
| 12 |
+
"transfer_v9_test": "3/20"
|
| 13 |
+
},
|
| 14 |
+
"accuracy_mean": "Exact weighted categorical accuracy; originalwithin-panelweights and requested denominators unchanged.",
|
| 15 |
+
"original_panel_registry_sha256": "1aefdf58d56e832e765194d0fa53a8647acea7d2ad2fc25ac95380bf85949fd2",
|
| 16 |
+
"transfer_protocol_sha256": "f0c0f8982bfbb9edfc5e486f682906612fae579f28c58c2fa656ad9d10874d16",
|
| 17 |
+
"transfer_payload_sha256": "97d82909bcbd8756904d6ed59b31e16d16086d62990ccdc981dc95b8c274e0c5",
|
| 18 |
+
"transfer_labels_sha256": "0b64882d58e2ac1e0f5e2e102f916907554c94c07f2774b8c092683cfc81973c",
|
| 19 |
+
"transfer_accuracy_questions": 1046,
|
| 20 |
+
"original_core_questions": 2720,
|
| 21 |
+
"diagnostics": "Preserve internal-v2 source-generalization/order/unknown/proper-score/efficiency axes separately; no millisecond or ECE averaging into accuracy.",
|
| 22 |
+
"roster": {
|
| 23 |
+
"Nox": {
|
| 24 |
+
"role": "Decision family",
|
| 25 |
+
"size": "4B",
|
| 26 |
+
"repo": "llm-semantic-router/Decision-1.0-Nox"
|
| 27 |
+
},
|
| 28 |
+
"Sol": {
|
| 29 |
+
"role": "Decision family",
|
| 30 |
+
"size": "2B",
|
| 31 |
+
"repo": "llm-semantic-router/Decision-1.0-Sol"
|
| 32 |
+
},
|
| 33 |
+
"Kai": {
|
| 34 |
+
"role": "Decision family",
|
| 35 |
+
"size": "0.572B encoder",
|
| 36 |
+
"repo": "llm-semantic-router/Decision-1.0-Kai",
|
| 37 |
+
"revision": "2079354070d02f4c2b1e60bce6b99f60dfad9622",
|
| 38 |
+
"contract": "shipped1024complete-token admission; FP32; defaultB8; no limit override"
|
| 39 |
+
},
|
| 40 |
+
"Lux": {
|
| 41 |
+
"role": "Decision family",
|
| 42 |
+
"size": "9B",
|
| 43 |
+
"repo": "llm-semantic-router/Decision-1.0-Lux",
|
| 44 |
+
"pending_unreleased": true
|
| 45 |
+
},
|
| 46 |
+
"kev-0.8b": {
|
| 47 |
+
"role": "open reference",
|
| 48 |
+
"size": "0.8B",
|
| 49 |
+
"repo": "jaredpalmer/kev-0.8b",
|
| 50 |
+
"revision": "54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8"
|
| 51 |
+
},
|
| 52 |
+
"kev-4b": {
|
| 53 |
+
"role": "open reference",
|
| 54 |
+
"size": "4B",
|
| 55 |
+
"repo": "jaredpalmer/kev-4b",
|
| 56 |
+
"revision": "485ace8703592fcf405488b262449990824cfed1"
|
| 57 |
+
},
|
| 58 |
+
"kev-9b": {
|
| 59 |
+
"role": "open reference",
|
| 60 |
+
"size": "9B",
|
| 61 |
+
"repo": "jaredpalmer/kev-9b",
|
| 62 |
+
"revision": "2629c06a5aeb0feb3b9783bafed17ed8f39ecf5c"
|
| 63 |
+
},
|
| 64 |
+
"Decider": {
|
| 65 |
+
"role": "open reference",
|
| 66 |
+
"size": "2B",
|
| 67 |
+
"repo": "Mapika/decider-2b"
|
| 68 |
+
},
|
| 69 |
+
"Laya-base": {
|
| 70 |
+
"role": "open reference",
|
| 71 |
+
"size": "0.421B encoder",
|
| 72 |
+
"repo": "convaiinnovations/laya",
|
| 73 |
+
"mode": "english"
|
| 74 |
+
},
|
| 75 |
+
"Laya-multilingual": {
|
| 76 |
+
"role": "open reference",
|
| 77 |
+
"size": "0.322B encoder",
|
| 78 |
+
"repo": "convaiinnovations/laya",
|
| 79 |
+
"mode": "multilingual"
|
| 80 |
+
},
|
| 81 |
+
"Jev": {
|
| 82 |
+
"role": "closed frontier",
|
| 83 |
+
"size": "undisclosed",
|
| 84 |
+
"served_identity": "jev-1.13.0"
|
| 85 |
+
},
|
| 86 |
+
"Qwen3.5-2B": {
|
| 87 |
+
"role": "untuned reference",
|
| 88 |
+
"size": "2B",
|
| 89 |
+
"repo": "Qwen/Qwen3.5-2B",
|
| 90 |
+
"adapter": "original locked chat/LM-head letter readout; no training"
|
| 91 |
+
},
|
| 92 |
+
"Qwen3.5-4B": {
|
| 93 |
+
"role": "untuned reference",
|
| 94 |
+
"size": "4B",
|
| 95 |
+
"repo": "Qwen/Qwen3.5-4B",
|
| 96 |
+
"adapter": "original locked chat/LM-head letter readout; no training"
|
| 97 |
+
},
|
| 98 |
+
"Qwen3.5-9B": {
|
| 99 |
+
"role": "untuned reference",
|
| 100 |
+
"size": "9B",
|
| 101 |
+
"repo": "Qwen/Qwen3.5-9B",
|
| 102 |
+
"adapter": "original locked chat/LM-head letter readout; no training"
|
| 103 |
+
}
|
| 104 |
+
},
|
| 105 |
+
"internal_extra_references": [
|
| 106 |
+
"LLM2Jev2B",
|
| 107 |
+
"LLM2Jev4B",
|
| 108 |
+
"Nimble9B"
|
| 109 |
+
],
|
| 110 |
+
"cohort_gates": {
|
| 111 |
+
"Sol": {
|
| 112 |
+
"required_same_size_open": [
|
| 113 |
+
"Decider"
|
| 114 |
+
],
|
| 115 |
+
"additional_measured_internal_peer": [
|
| 116 |
+
"LLM2Jev2B"
|
| 117 |
+
],
|
| 118 |
+
"required_same_size_untuned": [
|
| 119 |
+
"Qwen3.5-2B"
|
| 120 |
+
]
|
| 121 |
+
},
|
| 122 |
+
"Nox": {
|
| 123 |
+
"required_same_size_open": [
|
| 124 |
+
"kev-4b"
|
| 125 |
+
],
|
| 126 |
+
"additional_measured_internal_peer": [
|
| 127 |
+
"LLM2Jev4B"
|
| 128 |
+
],
|
| 129 |
+
"required_same_size_untuned": [
|
| 130 |
+
"Qwen3.5-4B"
|
| 131 |
+
]
|
| 132 |
+
},
|
| 133 |
+
"Lux": {
|
| 134 |
+
"must_exceed_current_family": [
|
| 135 |
+
"Nox",
|
| 136 |
+
"Sol"
|
| 137 |
+
],
|
| 138 |
+
"required_same_size_open": [
|
| 139 |
+
"kev-9b"
|
| 140 |
+
],
|
| 141 |
+
"additional_measured_internal_peer": [
|
| 142 |
+
"Nimble9B"
|
| 143 |
+
],
|
| 144 |
+
"required_same_size_untuned": [
|
| 145 |
+
"Qwen3.5-9B"
|
| 146 |
+
]
|
| 147 |
+
}
|
| 148 |
+
},
|
| 149 |
+
"publication": {
|
| 150 |
+
"first_expanded_cards": "Nox and Sol independently eligible only when their own required same-size comparisons are complete and strictly beaten; table includes complete available public roster. Unreleased Lux need not delay either card.",
|
| 151 |
+
"later_updates": "Require newaggregate strictly higher than the current published same model plus maintained same-size lead; show regressions and full diagnostic metrics honestly.",
|
| 152 |
+
"Lux_first_release": "Newweightedaggregate strictly exceeds then-currentNox/Sol and known same-sizeopen references; old4panelgate superseded before heldout/publication.",
|
| 153 |
+
"missing_scores": "Never fill from different benchmark, revision, successful-only denominator or inferred model size.",
|
| 154 |
+
"accuracy_uncertainty": "Report paired component bootstrap intervals; no new positive-CI gate invented.",
|
| 155 |
+
"versions": "No trainingversion labels in visiblecard/table/figures; retain immutable revision/weight/adapter identities in machine-readable provenance.",
|
| 156 |
+
"rank": "All public roster descending actual score. Decision family copper; distinct restrained colors for Kev, Decider, Laya, frontierJev and untunedQwen. White background, no logos or trainingversion labels.",
|
| 157 |
+
"matrix": "Minimal readable typography, full task coverage, no logo; labels omit trainingversions.",
|
| 158 |
+
"original_comparison": "Retain original four-panel and prior five-panel25/25/15/15/20 comparisons in methods/provenance. Explicit outcome-informed product-priority amendment; no universal or prospective performance claim."
|
| 159 |
+
},
|
| 160 |
+
"training_integrity": "Frozen ongoing training/SELECT/CAL unchanged. Testresults never select anothercheckpoint; future targetedtraining uses independent TRAIN/SELECT/CAL. Observed tests labeledregression.",
|
| 161 |
+
"supersedes": {
|
| 162 |
+
"protocol": {
|
| 163 |
+
"path": "eval/decision-benchmark-v3/PROTOCOL.json",
|
| 164 |
+
"sha256": "d0b7c667fd4de0dfa5268d2d83aebf40e2b0c43d615ad77f4c2264e35a137534"
|
| 165 |
+
},
|
| 166 |
+
"prior_full_statistics": {
|
| 167 |
+
"path": "analysis/decision-benchmark-v3-results/full01/scored/STATISTICS.json",
|
| 168 |
+
"sha256": "aaed8ca33a874300545b81566cb0ba37bf7cc8fcaac954beb9094aeef8ad82f7"
|
| 169 |
+
}
|
| 170 |
+
},
|
| 171 |
+
"sensitivity": "Retain full25/25/15/15/20 results alongside30/25/15/15/15 in methods/provenance. Reweighting gains are never described as training improvements. No model selection, calibration, predictions or denominators changed."
|
| 172 |
+
},
|
| 173 |
+
"statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
|
| 174 |
+
"manifest_sha256": "3199a40b2719cad9b1e815629cf5ebad923a9a399188c1199a9d27e918bda9e3",
|
| 175 |
+
"qualified_runtime": {
|
| 176 |
+
"python": "3.12.13",
|
| 177 |
+
"numpy": "2.3.5"
|
| 178 |
+
},
|
| 179 |
+
"source_sha256": "b00c60f21ace9bd258a1a79673b8cdbce3101ada2fed5547811f70bf030fcd91",
|
| 180 |
+
"model_manifest": {
|
| 181 |
+
"Nox": {
|
| 182 |
+
"panels": {
|
| 183 |
+
"old_core": {
|
| 184 |
+
"path": "results/expanded/nox-retention-v1/candidate/job-0/worker/old_core/normalized.jsonl",
|
| 185 |
+
"sha256": "9b6b7db907e8e458cd33982c0f7c45e5052c3248285d927990f160d5570c8976"
|
| 186 |
+
},
|
| 187 |
+
"v3_core": {
|
| 188 |
+
"path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v3_core/normalized.jsonl",
|
| 189 |
+
"sha256": "49575aef2261c191e289d781857664c57fd6181701c03f02515292f2fa31b40c"
|
| 190 |
+
},
|
| 191 |
+
"v4": {
|
| 192 |
+
"path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v4/normalized.jsonl",
|
| 193 |
+
"sha256": "9d85618b1a05c2ed68f6a188dadf19a64d6ede6d5a1b0ef58afa27a13906238b"
|
| 194 |
+
},
|
| 195 |
+
"v5": {
|
| 196 |
+
"path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v5/normalized.jsonl",
|
| 197 |
+
"sha256": "5eb7faaa45a22dad414af1363d6087b8c379c86599df3c14483a3a648ca6bfd5"
|
| 198 |
+
}
|
| 199 |
+
},
|
| 200 |
+
"transfer_rows": {
|
| 201 |
+
"path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Nox/v9-transfer-v9-test-published-temperature-rows.json",
|
| 202 |
+
"sha256": "77b92a18433b260ac0e6f4e0b6c9ec3a94e16aaed870a4078b1d7f32ef18b169"
|
| 203 |
+
},
|
| 204 |
+
"evidence": [
|
| 205 |
+
{
|
| 206 |
+
"path": "results/expanded/nox-retention-v1/QUALITY-COLLECTION.json",
|
| 207 |
+
"sha256": "7608258282ec1e9731e4d2ea4dee7206fda8900055dc64f9193d2566172ca86a"
|
| 208 |
+
},
|
| 209 |
+
{
|
| 210 |
+
"path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Nox/REPORT.json",
|
| 211 |
+
"sha256": "8e78910746c9e78862b3aadd94490054a681f3ad013b43ec19bfbc2434111ed7"
|
| 212 |
+
}
|
| 213 |
+
]
|
| 214 |
+
},
|
| 215 |
+
"Sol": {
|
| 216 |
+
"panels": {
|
| 217 |
+
"old_core": {
|
| 218 |
+
"path": "results/expanded/sol-composition-v2/candidate/job-0/worker/old_core/normalized.jsonl",
|
| 219 |
+
"sha256": "5fc0b3e49a5de576f7621839130855c178b12d942dc357aaf6b4dd9735e0a6a8"
|
| 220 |
+
},
|
| 221 |
+
"v3_core": {
|
| 222 |
+
"path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v3_core/normalized.jsonl",
|
| 223 |
+
"sha256": "b2c99fbd739b849a66a55929a658de4517a7ffaa4496255566e6972ba28d4f15"
|
| 224 |
+
},
|
| 225 |
+
"v4": {
|
| 226 |
+
"path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v4/normalized.jsonl",
|
| 227 |
+
"sha256": "3c28302750e02d4cead0aacd66a323734e773f726c4c41db658dd6461251b28b"
|
| 228 |
+
},
|
| 229 |
+
"v5": {
|
| 230 |
+
"path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v5/normalized.jsonl",
|
| 231 |
+
"sha256": "6ddb0496012b596ea122afa4db4c4158e57dfb934e9ae29eb006756b652e31e1"
|
| 232 |
+
}
|
| 233 |
+
},
|
| 234 |
+
"transfer_rows": {
|
| 235 |
+
"path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/v9-transfer-v9-test-published-temperature-rows.json",
|
| 236 |
+
"sha256": "b05799417cb3715de633135c58f713960b75bdb9ed30eb8912a197d33ee3d85b"
|
| 237 |
+
},
|
| 238 |
+
"evidence": [
|
| 239 |
+
{
|
| 240 |
+
"path": "results/expanded/sol-composition-v2/QUALITY-COLLECTION.json",
|
| 241 |
+
"sha256": "9b77aac87b85cbf3a74e86af4a2b9a03fd9e1ed854356a92ff54ccd0d47a16f4"
|
| 242 |
+
},
|
| 243 |
+
{
|
| 244 |
+
"path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/REPORT.json",
|
| 245 |
+
"sha256": "5d8ebb339293687cec9089f54fb96aabf3330f16111487a49489f5e52bbb6f29"
|
| 246 |
+
}
|
| 247 |
+
]
|
| 248 |
+
},
|
| 249 |
+
"kev-0.8b": {
|
| 250 |
+
"panels": {
|
| 251 |
+
"old_core": {
|
| 252 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-old_core-normalized.jsonl",
|
| 253 |
+
"sha256": "b492b4512ac164f83a319f31bbc9e086cae12f87ffe2c14f307481ffb240807c"
|
| 254 |
+
},
|
| 255 |
+
"v3_core": {
|
| 256 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v3_core-normalized.jsonl",
|
| 257 |
+
"sha256": "70cbfc69dfc72ac0c5ae01305ec63a9f31129c115b7068c6824cd22ead013d68"
|
| 258 |
+
},
|
| 259 |
+
"v4": {
|
| 260 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v4-normalized.jsonl",
|
| 261 |
+
"sha256": "761dfd69acbface101f2c21fd27d2db684c5c8b8cc425890a075729ae1f7db24"
|
| 262 |
+
},
|
| 263 |
+
"v5": {
|
| 264 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v5-normalized.jsonl",
|
| 265 |
+
"sha256": "51b6fc5f826ca8db44df43885f34d857ddd5b5a66453bbdc1a83fcf488928bdb"
|
| 266 |
+
}
|
| 267 |
+
},
|
| 268 |
+
"transfer_rows": {
|
| 269 |
+
"path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/v9-transfer-v9-test-published-temperature-rows.json",
|
| 270 |
+
"sha256": "d6abed88001aeac630219466f02c1aad89fcbd354b3ae533507f3c2f7fe42898"
|
| 271 |
+
},
|
| 272 |
+
"evidence": [
|
| 273 |
+
{
|
| 274 |
+
"path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/REPORT.json",
|
| 275 |
+
"sha256": "6760dcbaa7d85e6c2e893e4c80ab8e66e0b84d7cfa26ae9e25909f48a05f8de4"
|
| 276 |
+
},
|
| 277 |
+
{
|
| 278 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
|
| 279 |
+
"sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
|
| 280 |
+
}
|
| 281 |
+
]
|
| 282 |
+
},
|
| 283 |
+
"kev-4b": {
|
| 284 |
+
"panels": {
|
| 285 |
+
"old_core": {
|
| 286 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-old_core-normalized.jsonl",
|
| 287 |
+
"sha256": "41dfbb7c9c4381effc219f9123b51bd4d688b3e746cec974ef930730b9dfd6b1"
|
| 288 |
+
},
|
| 289 |
+
"v3_core": {
|
| 290 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v3_core-normalized.jsonl",
|
| 291 |
+
"sha256": "2ec54da92c4e72412e89b3c07013b54729f5899ba8542eeadf962b04a1f3a4b8"
|
| 292 |
+
},
|
| 293 |
+
"v4": {
|
| 294 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v4-normalized.jsonl",
|
| 295 |
+
"sha256": "88c576c3d62ec11686dd5ebf360971a9087779e45a346897e2badf874a754a12"
|
| 296 |
+
},
|
| 297 |
+
"v5": {
|
| 298 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v5-normalized.jsonl",
|
| 299 |
+
"sha256": "d33bbb60134314dcb550060176ed0a3ae91ec0c6b033cceb40e3255b33666898"
|
| 300 |
+
}
|
| 301 |
+
},
|
| 302 |
+
"transfer_rows": {
|
| 303 |
+
"path": "analysis/kev-reciprocal-v1/reports/kev-4b/v9-transfer-v9-test-published-temperature-rows.json",
|
| 304 |
+
"sha256": "c2cb16a936be4c10ca6b1870c69426ddd6b51629c401f71ae1da652b3bd2ac7e"
|
| 305 |
+
},
|
| 306 |
+
"evidence": [
|
| 307 |
+
{
|
| 308 |
+
"path": "analysis/kev-reciprocal-v1/reports/kev-4b/REPORT.json",
|
| 309 |
+
"sha256": "366fb0d27a4d732e5e270c264521d9a93ccabc4f0ddf5b7d3e9f1ecc5f460406"
|
| 310 |
+
},
|
| 311 |
+
{
|
| 312 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
|
| 313 |
+
"sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
|
| 314 |
+
}
|
| 315 |
+
]
|
| 316 |
+
},
|
| 317 |
+
"kev-9b": {
|
| 318 |
+
"panels": {
|
| 319 |
+
"old_core": {
|
| 320 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-old_core-normalized.jsonl",
|
| 321 |
+
"sha256": "a3b11fee4bc5e207cc42e7f4d8dde05df6cfc73d9446ac33cd6015e1ab9e587d"
|
| 322 |
+
},
|
| 323 |
+
"v3_core": {
|
| 324 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-v3_core-normalized.jsonl",
|
| 325 |
+
"sha256": "bb90e5701a151272f97caa4107888ba2379969b4e3f22abb4670a528b50d7729"
|
| 326 |
+
},
|
| 327 |
+
"v4": {
|
| 328 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-v4-normalized.jsonl",
|
| 329 |
+
"sha256": "b9c50612c0be1b7b5647779d8489c185ea7f02673da176b1d965673da604d644"
|
| 330 |
+
},
|
| 331 |
+
"v5": {
|
| 332 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-v5-normalized.jsonl",
|
| 333 |
+
"sha256": "2893714030357718e662633fddf31a328e56a02fc3ae71ed34612f480e50d4ee"
|
| 334 |
+
}
|
| 335 |
+
},
|
| 336 |
+
"transfer_rows": {
|
| 337 |
+
"path": "analysis/kev-reciprocal-v1/reports/kev-9b/v9-transfer-v9-test-published-temperature-rows.json",
|
| 338 |
+
"sha256": "e712bf23acea6a9431001f5424639d1e3738789907ff2e367db7c9dad4e0d239"
|
| 339 |
+
},
|
| 340 |
+
"evidence": [
|
| 341 |
+
{
|
| 342 |
+
"path": "analysis/kev-reciprocal-v1/reports/kev-9b/REPORT.json",
|
| 343 |
+
"sha256": "96dcf11961f510d05c011aa3fd8dfa6043348471b8e800d6bac2d691d0815d54"
|
| 344 |
+
},
|
| 345 |
+
{
|
| 346 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
|
| 347 |
+
"sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
|
| 348 |
+
}
|
| 349 |
+
]
|
| 350 |
+
},
|
| 351 |
+
"Jev": {
|
| 352 |
+
"panels": {
|
| 353 |
+
"old_core": {
|
| 354 |
+
"path": "eval/heldout/official-predictions.jsonl",
|
| 355 |
+
"sha256": "069e48e1fdf1b5f3d95cc0097e4179ac3faab97bb357b0ce897cfb463b87cbaf",
|
| 356 |
+
"bytes": 185263
|
| 357 |
+
},
|
| 358 |
+
"v3_core": {
|
| 359 |
+
"path": "eval/heldout/v3/official/normalized/core.jsonl",
|
| 360 |
+
"sha256": "54707354bfe5d96b5de56e6dc1b38cba0a610eea516b28fbff013b94f422b397",
|
| 361 |
+
"bytes": 368716
|
| 362 |
+
},
|
| 363 |
+
"v4": {
|
| 364 |
+
"path": "eval/heldout/v4/official/normalized/core.jsonl",
|
| 365 |
+
"sha256": "d6583f903e974d70747fae2944014828954dc4d7344ceddcbb2cd1b06fbfc799",
|
| 366 |
+
"bytes": 183671
|
| 367 |
+
},
|
| 368 |
+
"v5": {
|
| 369 |
+
"path": "eval/heldout/v5/official/normalized/core.jsonl",
|
| 370 |
+
"sha256": "f883164c4c83db8e42ebe03f7e165afbd2f81f2b541bcf8008a2cc06a60e1ee7",
|
| 371 |
+
"bytes": 200316
|
| 372 |
+
}
|
| 373 |
+
},
|
| 374 |
+
"transfer_rows": {
|
| 375 |
+
"path": "analysis/decision-benchmark-v3-official/scored/ROWS.json",
|
| 376 |
+
"sha256": "be29a7f3804ad84993ffb06f05ada7e2d8ed25317504c4fc17fe5dc306fd788c"
|
| 377 |
+
},
|
| 378 |
+
"evidence": [
|
| 379 |
+
{
|
| 380 |
+
"path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
|
| 381 |
+
"sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
|
| 382 |
+
},
|
| 383 |
+
{
|
| 384 |
+
"path": "analysis/decision-benchmark-v3-official/scored/REPORT.json",
|
| 385 |
+
"sha256": "2fdbaf7db1e4b877762e33a87b23a236a065eaa3078b14c2d7d579ad44f9ea24"
|
| 386 |
+
},
|
| 387 |
+
{
|
| 388 |
+
"path": "analysis/decision-benchmark-v3-official/INDEPENDENT-PROJECTION-REVIEW.json",
|
| 389 |
+
"sha256": "b9506539e2c4a863aef16f7dab0517c175e85863416bc622c60df05489a9eb52"
|
| 390 |
+
}
|
| 391 |
+
]
|
| 392 |
+
},
|
| 393 |
+
"Kai": {
|
| 394 |
+
"panels": {
|
| 395 |
+
"old_core": {
|
| 396 |
+
"path": "results/public-roster-baselines-v1/Kai/worker/old_core/normalized.jsonl",
|
| 397 |
+
"sha256": "4758a99a75392b197c8902600a88edca65d6acea500e046e1ea8e4edd398a8a3"
|
| 398 |
+
},
|
| 399 |
+
"v3_core": {
|
| 400 |
+
"path": "results/public-roster-baselines-v1/Kai/worker/v3_core/normalized.jsonl",
|
| 401 |
+
"sha256": "7cdb285b88882a4c0d73c3da890056ef3b1d09e09bc01a893e73a12ff3bdbab8"
|
| 402 |
+
},
|
| 403 |
+
"v4": {
|
| 404 |
+
"path": "results/public-roster-baselines-v1/Kai/worker/v4/normalized.jsonl",
|
| 405 |
+
"sha256": "fa744ee2181fc4c379b748fdd4ee1080ffd877f43ebc9441d27a47013a2a9cec"
|
| 406 |
+
},
|
| 407 |
+
"v5": {
|
| 408 |
+
"path": "results/public-roster-baselines-v1/Kai/worker/v5/normalized.jsonl",
|
| 409 |
+
"sha256": "9d1e50d62733dbf06b824855938a9ee1311c030916cf173cc5b18fe066af48cd"
|
| 410 |
+
}
|
| 411 |
+
},
|
| 412 |
+
"transfer_rows": {
|
| 413 |
+
"path": "analysis/decision-benchmark-v3-baselines/Kai/ROWS.json",
|
| 414 |
+
"sha256": "854950cb3888a49a801982341576f4a2ebdacbdd7e7a7e32d16f5777e0cbda46"
|
| 415 |
+
},
|
| 416 |
+
"evidence": [
|
| 417 |
+
{
|
| 418 |
+
"path": "results/public-roster-baselines-v1/Kai/worker/COMPLETE.json",
|
| 419 |
+
"sha256": "39cc8bae7b59e02720123cd8efa9ba1939e02ef15a4bb3be89972b14eb7877cf"
|
| 420 |
+
},
|
| 421 |
+
{
|
| 422 |
+
"path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
|
| 423 |
+
"sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
|
| 424 |
+
},
|
| 425 |
+
{
|
| 426 |
+
"path": "analysis/decision-benchmark-v3-baselines/Kai/REPORT.json",
|
| 427 |
+
"sha256": "5ed1fd6a69d68fbd64bd45c4da9431a73fd6053574786ddd00323e0537fcefbc"
|
| 428 |
+
}
|
| 429 |
+
]
|
| 430 |
+
},
|
| 431 |
+
"Laya-base": {
|
| 432 |
+
"panels": {
|
| 433 |
+
"v4": {
|
| 434 |
+
"path": "results/public-roster-baselines-v1/laya-english/worker/v4/normalized.jsonl",
|
| 435 |
+
"sha256": "31da1d4140da69d2677172707eb301843cf8fdfb43200cba93be910321f922ba"
|
| 436 |
+
},
|
| 437 |
+
"v5": {
|
| 438 |
+
"path": "results/public-roster-baselines-v1/laya-english/worker/v5/normalized.jsonl",
|
| 439 |
+
"sha256": "159614ceede814b578e1e025836a595c8c71a9bec4896caf6454d7bf54c4c44d"
|
| 440 |
+
},
|
| 441 |
+
"old_core": {
|
| 442 |
+
"path": "eval/heldout/primary-v2/laya-english/core/normalized.jsonl",
|
| 443 |
+
"sha256": "850643d2daeb47a6909985003736cab0f1f0e5a87631a35e27192f0733efd756"
|
| 444 |
+
},
|
| 445 |
+
"v3_core": {
|
| 446 |
+
"path": "eval/heldout/v3/baselines/laya-english/core/normalized.jsonl",
|
| 447 |
+
"sha256": "dec19caeb8074e25289e43a367392172da21f0c2ccdb403a059e2e4a2364a64b"
|
| 448 |
+
}
|
| 449 |
+
},
|
| 450 |
+
"transfer_rows": {
|
| 451 |
+
"path": "analysis/decision-benchmark-v3-baselines/laya-english/ROWS.json",
|
| 452 |
+
"sha256": "19138909259400facb5d3d5402d0d736f870695fe3ee46019f2e438088c57275"
|
| 453 |
+
},
|
| 454 |
+
"evidence": [
|
| 455 |
+
{
|
| 456 |
+
"path": "results/public-roster-baselines-v1/laya-english/worker/COMPLETE.json",
|
| 457 |
+
"sha256": "cef0da797ae14cebaf1e4b9edf9b16c3cd9a412a30435d7d2517b8ba6dbe8b98"
|
| 458 |
+
},
|
| 459 |
+
{
|
| 460 |
+
"path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
|
| 461 |
+
"sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
|
| 462 |
+
},
|
| 463 |
+
{
|
| 464 |
+
"path": "analysis/decision-benchmark-v3-baselines/laya-english/REPORT.json",
|
| 465 |
+
"sha256": "cb5935bb399dcf23ffe24f855452b2e074bcbb3d023cfc7dd439342e54f02c6b"
|
| 466 |
+
},
|
| 467 |
+
{
|
| 468 |
+
"path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
|
| 469 |
+
"sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
|
| 470 |
+
}
|
| 471 |
+
]
|
| 472 |
+
},
|
| 473 |
+
"Laya-multilingual": {
|
| 474 |
+
"panels": {
|
| 475 |
+
"v4": {
|
| 476 |
+
"path": "results/public-roster-baselines-v1/laya-multilingual/worker/v4/normalized.jsonl",
|
| 477 |
+
"sha256": "814c1aa1f4d400942984ecb3786d3fd5fd15286d9b767a519c91993577795290"
|
| 478 |
+
},
|
| 479 |
+
"v5": {
|
| 480 |
+
"path": "results/public-roster-baselines-v1/laya-multilingual/worker/v5/normalized.jsonl",
|
| 481 |
+
"sha256": "6d4dbc16c7b977b15895e0e313dabc6d9db43df75ec56e89373d91f4f13ec676"
|
| 482 |
+
},
|
| 483 |
+
"old_core": {
|
| 484 |
+
"path": "eval/heldout/primary-v2/laya-multilingual/core/normalized.jsonl",
|
| 485 |
+
"sha256": "9ea39841c4d5901cd9ab860c980d2f3b55358be289c7e01b9b47982927bf849d"
|
| 486 |
+
},
|
| 487 |
+
"v3_core": {
|
| 488 |
+
"path": "eval/heldout/v3/baselines/laya-multilingual/core/normalized.jsonl",
|
| 489 |
+
"sha256": "a0b7ce9daf4be2c7831bcab6b5e96c8dcb6bef133d39475f72241ab18faccdd2"
|
| 490 |
+
}
|
| 491 |
+
},
|
| 492 |
+
"transfer_rows": {
|
| 493 |
+
"path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/ROWS.json",
|
| 494 |
+
"sha256": "d32db6cf6576f287210d31336456497f9d6926179366db29304659baf9783b43"
|
| 495 |
+
},
|
| 496 |
+
"evidence": [
|
| 497 |
+
{
|
| 498 |
+
"path": "results/public-roster-baselines-v1/laya-multilingual/worker/COMPLETE.json",
|
| 499 |
+
"sha256": "b190d683fc4c4762c2805a03d52628a53e125d228a3105c8ac2ddda00a5364e4"
|
| 500 |
+
},
|
| 501 |
+
{
|
| 502 |
+
"path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
|
| 503 |
+
"sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
|
| 504 |
+
},
|
| 505 |
+
{
|
| 506 |
+
"path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/REPORT.json",
|
| 507 |
+
"sha256": "d20d3d58f76ca2758545f83d358abd747f86d581c897d1fb521371d1000d6154"
|
| 508 |
+
},
|
| 509 |
+
{
|
| 510 |
+
"path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
|
| 511 |
+
"sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
|
| 512 |
+
}
|
| 513 |
+
]
|
| 514 |
+
},
|
| 515 |
+
"Qwen3.5-9B": {
|
| 516 |
+
"panels": {
|
| 517 |
+
"old_core": {
|
| 518 |
+
"path": "results/lux-new-host-v1/quality/base9b/old_core/normalized.jsonl",
|
| 519 |
+
"sha256": "16ccd1c30bbf29560a7ce1a288143692c5d2ed69cc7afebdc146d5bcc4b77611"
|
| 520 |
+
},
|
| 521 |
+
"v3_core": {
|
| 522 |
+
"path": "results/lux-new-host-v1/quality/base9b/v3_core/normalized.jsonl",
|
| 523 |
+
"sha256": "d081a8158863e85d5e49cb853fee87192119767124f53f9e6c782392f8edb018"
|
| 524 |
+
},
|
| 525 |
+
"v4": {
|
| 526 |
+
"path": "results/lux-new-host-v1/quality/base9b/v4/normalized.jsonl",
|
| 527 |
+
"sha256": "cc2775884d84c49e5155afbebc7f4562ec86f6f6d32765ffcbdafe653ca7d132"
|
| 528 |
+
},
|
| 529 |
+
"v5": {
|
| 530 |
+
"path": "results/lux-new-host-v1/quality/base9b/v5/normalized.jsonl",
|
| 531 |
+
"sha256": "ce6bf5c286e5920055d7de36fbd910311f7d3185380c63681c7baea6d0e3494a"
|
| 532 |
+
}
|
| 533 |
+
},
|
| 534 |
+
"transfer_rows": {
|
| 535 |
+
"path": "analysis/decision-benchmark-v3-baselines/Qwen3.5-9B/ROWS.json",
|
| 536 |
+
"sha256": "5acdd41f616a3ba63a5407af9c8371268012792fd324b70230d899bbb37101b4"
|
| 537 |
+
},
|
| 538 |
+
"evidence": [
|
| 539 |
+
{
|
| 540 |
+
"path": "results/lux-new-host-v1/quality/base9b/COMPLETE.json",
|
| 541 |
+
"sha256": "702c65f5c290f5cdda0495479f361101dbe976330233c5359455ba736245c3aa"
|
| 542 |
+
},
|
| 543 |
+
{
|
| 544 |
+
"path": "results/lux-new-host-v1/quality/base9b/metadata.json",
|
| 545 |
+
"sha256": "eaa29f71aba60e86068e6c5d1b5782ea73b0214b81a297b57419c4f0a2fc6c04"
|
| 546 |
+
},
|
| 547 |
+
{
|
| 548 |
+
"path": "results/lux-new-host-v1/quality/base9b/old_core/complete.json",
|
| 549 |
+
"sha256": "d28f2a646a908fc775e3727569a608f30ae5820f88a6173750a1031604e84f47"
|
| 550 |
+
},
|
| 551 |
+
{
|
| 552 |
+
"path": "results/lux-new-host-v1/quality/base9b/v3_core/complete.json",
|
| 553 |
+
"sha256": "3a10bd5e2121a81f8ec6f76b2526723a17689ec34d13f062c5259983203992cc"
|
| 554 |
+
},
|
| 555 |
+
{
|
| 556 |
+
"path": "results/lux-new-host-v1/quality/base9b/v4/complete.json",
|
| 557 |
+
"sha256": "073239ab7dd21768bccdfa4aaf3217c99dc7ff9b00fc308ffbe889c72fe002b4"
|
| 558 |
+
},
|
| 559 |
+
{
|
| 560 |
+
"path": "results/lux-new-host-v1/quality/base9b/v5/complete.json",
|
| 561 |
+
"sha256": "9a8d1e8e8db066f1f6594f137a7a1571dd5184aae363a9f919e44b634b782bba"
|
| 562 |
+
},
|
| 563 |
+
{
|
| 564 |
+
"path": "analysis/decision-benchmark-v3-baselines/Qwen3.5-9B/REPORT.json",
|
| 565 |
+
"sha256": "31b0a92bbbe7a0d259cb8e296a10b04b2bdd685b394eabdc5fbf6da1556ff12d"
|
| 566 |
+
},
|
| 567 |
+
{
|
| 568 |
+
"path": "analysis/decision-benchmark-v3-baselines/Qwen3.5-9B-projected/COMPLETE.json",
|
| 569 |
+
"sha256": "eaff79adb3a6a6bad61253a01b4c4e22cfa3d9969cd963cd82a3685755b77ec1"
|
| 570 |
+
}
|
| 571 |
+
]
|
| 572 |
+
},
|
| 573 |
+
"Lux": {
|
| 574 |
+
"panels": {
|
| 575 |
+
"old_core": {
|
| 576 |
+
"path": "results/lux-new-host-v1/quality/lux/old_core/normalized.jsonl",
|
| 577 |
+
"sha256": "a659a7c7fef2871bbf806e842ea8ca81f645dcf610f0f2f6622d8d32e76a0836"
|
| 578 |
+
},
|
| 579 |
+
"v3_core": {
|
| 580 |
+
"path": "results/lux-new-host-v1/quality/lux/v3_core/normalized.jsonl",
|
| 581 |
+
"sha256": "da8903a07e36a9687f6ee1d0fd8dbc684e61c49ef5fea30c346e46bff5a90e94"
|
| 582 |
+
},
|
| 583 |
+
"v4": {
|
| 584 |
+
"path": "results/lux-new-host-v1/quality/lux/v4/normalized.jsonl",
|
| 585 |
+
"sha256": "26cd7482693ec1ffa1a93a156abfc88378e4872e320f3b9b775763635086d488"
|
| 586 |
+
},
|
| 587 |
+
"v5": {
|
| 588 |
+
"path": "results/lux-new-host-v1/quality/lux/v5/normalized.jsonl",
|
| 589 |
+
"sha256": "759c773e629a16317271de850780a8734533c795046122f850914c4482d1e72b"
|
| 590 |
+
}
|
| 591 |
+
},
|
| 592 |
+
"transfer_rows": {
|
| 593 |
+
"path": "analysis/decision-benchmark-v3-baselines/Lux/ROWS.json",
|
| 594 |
+
"sha256": "bafb55fc86762061e537382360eda14a8a8f8c5e4e111c2fda1f82cd6c06793f"
|
| 595 |
+
},
|
| 596 |
+
"evidence": [
|
| 597 |
+
{
|
| 598 |
+
"path": "results/lux-new-host-v1/quality/lux/COMPLETE.json",
|
| 599 |
+
"sha256": "9600ff792cbda13af5505c804cf63af94d2ee895f2912f0645a7d493ae1b7a9c"
|
| 600 |
+
},
|
| 601 |
+
{
|
| 602 |
+
"path": "results/lux-new-host-v1/quality/lux/metadata.json",
|
| 603 |
+
"sha256": "0d4509fae7eb3630118eaebc24e75a7b02f717d9a0ad74811fa27358c33c9cc5"
|
| 604 |
+
},
|
| 605 |
+
{
|
| 606 |
+
"path": "analysis/decision-benchmark-v3-baselines/Lux-projected/COMPLETE.json",
|
| 607 |
+
"sha256": "f58d02df78f297d4feebe39facfac072951111d8d6a19e2c1069f1b90250c622"
|
| 608 |
+
},
|
| 609 |
+
{
|
| 610 |
+
"path": "analysis/decision-benchmark-v3-baselines/Lux/REPORT.json",
|
| 611 |
+
"sha256": "f31703c342efe043a1e1534da189808bb7e0de0641fbaf3f91faa4a1f928a299"
|
| 612 |
+
},
|
| 613 |
+
{
|
| 614 |
+
"path": "analysis/decoder4b/lux9b-training-v1/CHECKPOINT-HELDOUT-COLLECTED.json",
|
| 615 |
+
"sha256": "d1c42515ab225b27ea011a6cf2262fe12dd7001e3b4d2cb64a23e7642c7a0d91"
|
| 616 |
+
}
|
| 617 |
+
]
|
| 618 |
+
},
|
| 619 |
+
"Decider": {
|
| 620 |
+
"panels": {
|
| 621 |
+
"old_core": {
|
| 622 |
+
"path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/decider/quality/normalized.jsonl",
|
| 623 |
+
"sha256": "e1b14f0f9ec521bffbde5945684bc2410a413b6b6418d65ac4fc702a7b5e7937"
|
| 624 |
+
},
|
| 625 |
+
"v3_core": {
|
| 626 |
+
"path": "eval/heldout/v3/baselines/decider/core/normalized.jsonl",
|
| 627 |
+
"sha256": "12a1f554decf1aff608743e7b4a44681289ea084ae25f383e519b1f34917d46f"
|
| 628 |
+
},
|
| 629 |
+
"v4": {
|
| 630 |
+
"path": "eval/heldout/v4/baselines/decider/core/normalized.jsonl",
|
| 631 |
+
"sha256": "626ef366ae471ec10dfb89ef2ff2f7b29a6c879976d3ff14fcbb2bafd7d9b045"
|
| 632 |
+
},
|
| 633 |
+
"v5": {
|
| 634 |
+
"path": "eval/heldout/v5/open-baselines-v1/decider/core/normalized.jsonl",
|
| 635 |
+
"sha256": "ae16b17ac8e4208940c0c3e9e04f25e94686b5d16bf46f64f53669baf41aeb71"
|
| 636 |
+
}
|
| 637 |
+
},
|
| 638 |
+
"transfer_rows": {
|
| 639 |
+
"path": "analysis/decision-benchmark-v3-baselines/decider/ROWS.json",
|
| 640 |
+
"sha256": "6da5b5692a8a0cea3dc8f6d02d911fafd9bbf62e286a9be8db27c89f77a897e4"
|
| 641 |
+
},
|
| 642 |
+
"evidence": [
|
| 643 |
+
{
|
| 644 |
+
"path": "results/accelerated-transfer-baselines-v1/decider/worker/COMPLETE.json",
|
| 645 |
+
"sha256": "e1ee0df8df59ec3eb2b49ba947a8508994a24bce226b7f2cccbe25f6835b99af"
|
| 646 |
+
},
|
| 647 |
+
{
|
| 648 |
+
"path": "results/accelerated-transfer-baselines-v1/decider/worker/METADATA.json",
|
| 649 |
+
"sha256": "f00061c1b0bb642bf1353aae867d19df988c01b0428b58206b31249521bcbff8"
|
| 650 |
+
},
|
| 651 |
+
{
|
| 652 |
+
"path": "results/accelerated-transfer-baselines-v1/decider/worker/predictions.jsonl",
|
| 653 |
+
"sha256": "f69e87fa65b09289e6861c1d789270aea5a92d23ac47817342647ed08f406df7"
|
| 654 |
+
},
|
| 655 |
+
{
|
| 656 |
+
"path": "analysis/decision-benchmark-v3-baselines/decider/REPORT.json",
|
| 657 |
+
"sha256": "5aae029f34524dfebc8e79128da8c39229c5747461dfce07ca59956da4416951"
|
| 658 |
+
},
|
| 659 |
+
{
|
| 660 |
+
"path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
|
| 661 |
+
"sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
|
| 662 |
+
},
|
| 663 |
+
{
|
| 664 |
+
"path": "analysis/decision-benchmark-v3-baselines/decider/COMPOSABLE-MANIFEST.json",
|
| 665 |
+
"sha256": "a49848e86feafaa4f36a33287910035e7605e344791aae139305b78cb304a77f"
|
| 666 |
+
}
|
| 667 |
+
]
|
| 668 |
+
},
|
| 669 |
+
"Qwen3.5-2B": {
|
| 670 |
+
"panels": {
|
| 671 |
+
"old_core": {
|
| 672 |
+
"path": "eval/heldout/primary-v2/base-2b/core/normalized.jsonl",
|
| 673 |
+
"sha256": "c11a94d8f833d83427ea984a4f069aefaa3fa10a2b1085fb3a03542727f819cf"
|
| 674 |
+
},
|
| 675 |
+
"v3_core": {
|
| 676 |
+
"path": "eval/heldout/v3/baselines/base-2b/core/normalized.jsonl",
|
| 677 |
+
"sha256": "cc51471716342ecd466b50d5dd9b4ba805f5a6873cd96075b56d7c2e8336175e"
|
| 678 |
+
},
|
| 679 |
+
"v4": {
|
| 680 |
+
"path": "eval/heldout/v4/baselines/base-2b/core/normalized.jsonl",
|
| 681 |
+
"sha256": "93d84cdbe607691e2cef35ece3aa2dd5900f6410b0d72796b86d74c9156f3cdd"
|
| 682 |
+
},
|
| 683 |
+
"v5": {
|
| 684 |
+
"path": "eval/heldout/v5/open-baselines-v1/base-2b/core/normalized.jsonl",
|
| 685 |
+
"sha256": "75740c4279eb8dcc9c90b55f155916b3748c887e8c9d9e1d3a47edf7f4dbc34b"
|
| 686 |
+
}
|
| 687 |
+
},
|
| 688 |
+
"transfer_rows": {
|
| 689 |
+
"path": "analysis/decision-benchmark-v3-baselines/base-2b/ROWS.json",
|
| 690 |
+
"sha256": "62bf4f977d03eabd4f051aca5c8f7204f350eb995545aedeefa7bfa00bcbd115"
|
| 691 |
+
},
|
| 692 |
+
"evidence": [
|
| 693 |
+
{
|
| 694 |
+
"path": "results/accelerated-transfer-baselines-v1/base-2b/worker/COMPLETE.json",
|
| 695 |
+
"sha256": "0353d786b760d491c2a5ff3f8363ecd27a537f25a42dbeb2cea4db5d5eb77910"
|
| 696 |
+
},
|
| 697 |
+
{
|
| 698 |
+
"path": "results/accelerated-transfer-baselines-v1/base-2b/worker/METADATA.json",
|
| 699 |
+
"sha256": "745151846ade07431bc8ac964894bbd6ed234fb6a4665fcce223d91e09acaeeb"
|
| 700 |
+
},
|
| 701 |
+
{
|
| 702 |
+
"path": "results/accelerated-transfer-baselines-v1/base-2b/worker/predictions.jsonl",
|
| 703 |
+
"sha256": "347fc785ee59d5ca7694357b68a6c2e7bc7f863f1e26d94fc9b989391ed7d5fe"
|
| 704 |
+
},
|
| 705 |
+
{
|
| 706 |
+
"path": "analysis/decision-benchmark-v3-baselines/base-2b/REPORT.json",
|
| 707 |
+
"sha256": "9eb1db12e826a75dcda322960478bb153c4c1c24e5f5853d29a761b82bea36e7"
|
| 708 |
+
},
|
| 709 |
+
{
|
| 710 |
+
"path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
|
| 711 |
+
"sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
|
| 712 |
+
},
|
| 713 |
+
{
|
| 714 |
+
"path": "analysis/decision-benchmark-v3-baselines/base-2b/COMPOSABLE-MANIFEST.json",
|
| 715 |
+
"sha256": "c3e6f28b600776686afc56c1386cabf0869e69d5ca79f3ae6a0b17189eee968e"
|
| 716 |
+
}
|
| 717 |
+
]
|
| 718 |
+
},
|
| 719 |
+
"Qwen3.5-4B": {
|
| 720 |
+
"panels": {
|
| 721 |
+
"old_core": {
|
| 722 |
+
"path": "eval/heldout/primary-v2/base-4b/core/normalized.jsonl",
|
| 723 |
+
"sha256": "d20ad845066ff81caaf60ebee44df941872bdc642da78b5f3b3da4a33dd9a08d"
|
| 724 |
+
},
|
| 725 |
+
"v3_core": {
|
| 726 |
+
"path": "eval/heldout/v3/baselines/base-4b/core/normalized.jsonl",
|
| 727 |
+
"sha256": "7541893d2b2ec0ec2ffc513f50c568eaef860b65bcef9a1747996c116fff9661"
|
| 728 |
+
},
|
| 729 |
+
"v4": {
|
| 730 |
+
"path": "eval/heldout/v4/baselines/base-4b/core/normalized.jsonl",
|
| 731 |
+
"sha256": "7a6d252ccaa72bf54b94445190bbad023fb465a3eaf4cf657532e2cdd4f750f5"
|
| 732 |
+
},
|
| 733 |
+
"v5": {
|
| 734 |
+
"path": "eval/heldout/v5/open-baselines-v1/base-4b/core/normalized.jsonl",
|
| 735 |
+
"sha256": "c01ad5daed65267e9aec40ed9e97ece1b6511e4cab17ebff09339b41524fb7f1"
|
| 736 |
+
}
|
| 737 |
+
},
|
| 738 |
+
"transfer_rows": {
|
| 739 |
+
"path": "analysis/decision-benchmark-v3-baselines/base-4b/ROWS.json",
|
| 740 |
+
"sha256": "37ab17d94d640ee46283e1f23723b9ea05133c22a8a9b62a97c756a89c5de785"
|
| 741 |
+
},
|
| 742 |
+
"evidence": [
|
| 743 |
+
{
|
| 744 |
+
"path": "results/remaining-transfer-baselines-v1/base-4b/worker/COMPLETE.json",
|
| 745 |
+
"sha256": "72abc4f87a9234f95eb3b0a72a9e7d66798779d5982c757faffb610003b1b242"
|
| 746 |
+
},
|
| 747 |
+
{
|
| 748 |
+
"path": "results/remaining-transfer-baselines-v1/base-4b/worker/METADATA.json",
|
| 749 |
+
"sha256": "871fed92cf762187015e59174a042808dbf3fc88f7791aa0ddd530ae923196ab"
|
| 750 |
+
},
|
| 751 |
+
{
|
| 752 |
+
"path": "results/remaining-transfer-baselines-v1/base-4b/worker/predictions.jsonl",
|
| 753 |
+
"sha256": "8c47edda469f8914eccc2ecafa91c245752812e830cb9f29445653102522854a"
|
| 754 |
+
},
|
| 755 |
+
{
|
| 756 |
+
"path": "analysis/decision-benchmark-v3-baselines/base-4b/REPORT.json",
|
| 757 |
+
"sha256": "c2a0aafbe2b7c85b3c07a4b9843182a3361654f57883311155aa3b05cfc7002d"
|
| 758 |
+
},
|
| 759 |
+
{
|
| 760 |
+
"path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
|
| 761 |
+
"sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
|
| 762 |
+
},
|
| 763 |
+
{
|
| 764 |
+
"path": "analysis/decision-benchmark-v3-baselines/base-4b/COMPOSABLE-MANIFEST.json",
|
| 765 |
+
"sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
|
| 766 |
+
}
|
| 767 |
+
]
|
| 768 |
+
},
|
| 769 |
+
"llm2jev-2b": {
|
| 770 |
+
"panels": {
|
| 771 |
+
"old_core": {
|
| 772 |
+
"path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard0/llm2jev-2b/quality/normalized.jsonl",
|
| 773 |
+
"sha256": "5f0aaf2be47d11c42ecb228aeb6f317bf2074c80d398a66969f06490d5847760"
|
| 774 |
+
},
|
| 775 |
+
"v3_core": {
|
| 776 |
+
"path": "eval/heldout/v3/baselines/llm2jev-2b/core/normalized.jsonl",
|
| 777 |
+
"sha256": "339cd97706404b3d1e5af9d07ae6d3efc79cbb40310cfcf1406c64477e24e460"
|
| 778 |
+
},
|
| 779 |
+
"v4": {
|
| 780 |
+
"path": "results/open-reference-gap-v1/llm2jev-2b/v4/normalized.jsonl",
|
| 781 |
+
"sha256": "98e6b3bc2113e18e00442a86b085e06981560f43cf6a8807c3d1818b3fd27caf"
|
| 782 |
+
},
|
| 783 |
+
"v5": {
|
| 784 |
+
"path": "results/open-reference-gap-v1/llm2jev-2b/v5/normalized.jsonl",
|
| 785 |
+
"sha256": "36b9d9e96bae101eb3db6600c412585f48551ff5908072963a5dd435632e4d65"
|
| 786 |
+
}
|
| 787 |
+
},
|
| 788 |
+
"transfer_rows": {
|
| 789 |
+
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/ROWS.json",
|
| 790 |
+
"sha256": "a37a6b19549896854fd0fb9bb32f594bb8cb60b840aa49ead0c004dcea00778e"
|
| 791 |
+
},
|
| 792 |
+
"evidence": [
|
| 793 |
+
{
|
| 794 |
+
"path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/COMPLETE.json",
|
| 795 |
+
"sha256": "ccd7e1f0d6a4193c629046961eb7171b1278a8dbae015ff216396ceac1f008ef"
|
| 796 |
+
},
|
| 797 |
+
{
|
| 798 |
+
"path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/METADATA.json",
|
| 799 |
+
"sha256": "b0d0aa9840976e0330e676fdd67ce90230dbfec7502452956260188924583b2b"
|
| 800 |
+
},
|
| 801 |
+
{
|
| 802 |
+
"path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/predictions.jsonl",
|
| 803 |
+
"sha256": "67a217eea9af720846f653d780973a64fff93111d8dcb9fa09b8b6f9594328a4"
|
| 804 |
+
},
|
| 805 |
+
{
|
| 806 |
+
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/REPORT.json",
|
| 807 |
+
"sha256": "bb2bf635114fc805e6fef66dce70ac6ceca126008b3e8365871a6d12d9cd6e09"
|
| 808 |
+
},
|
| 809 |
+
{
|
| 810 |
+
"path": "eval/open-reference-comparison-v1/REGISTRY.json",
|
| 811 |
+
"sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
|
| 812 |
+
},
|
| 813 |
+
{
|
| 814 |
+
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/COMPOSABLE-MANIFEST.json",
|
| 815 |
+
"sha256": "071d3b47bd6a81e256e11be2ed2965c7416244bd512bd35be8baef5b8ebcb2be"
|
| 816 |
+
}
|
| 817 |
+
]
|
| 818 |
+
},
|
| 819 |
+
"llm2jev-4b": {
|
| 820 |
+
"panels": {
|
| 821 |
+
"old_core": {
|
| 822 |
+
"path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/llm2jev-4b/quality/normalized.jsonl",
|
| 823 |
+
"sha256": "8545cdc4ab9e9f870fdce1c1859632b8bac7505f6e961b4cc07e4f3b816328ce"
|
| 824 |
+
},
|
| 825 |
+
"v3_core": {
|
| 826 |
+
"path": "eval/heldout/v3/baselines/llm2jev-4b/core/normalized.jsonl",
|
| 827 |
+
"sha256": "b3bd38c90daf0bc3d77fea9a3ed0612049b07865bde5102356f82f7456002c97"
|
| 828 |
+
},
|
| 829 |
+
"v4": {
|
| 830 |
+
"path": "results/open-reference-gap-v1/llm2jev-4b/v4/normalized.jsonl",
|
| 831 |
+
"sha256": "409f8d2c57a1f8a4ef7dd18f3adb6cb8ee2ed544566206ccc909d54236ca95b7"
|
| 832 |
+
},
|
| 833 |
+
"v5": {
|
| 834 |
+
"path": "results/open-reference-gap-v1/llm2jev-4b/v5/normalized.jsonl",
|
| 835 |
+
"sha256": "d458e7e11cd2b1d66c7db6044408042630a48a474c9d3997141fc85dfb427f24"
|
| 836 |
+
}
|
| 837 |
+
},
|
| 838 |
+
"transfer_rows": {
|
| 839 |
+
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/ROWS.json",
|
| 840 |
+
"sha256": "abc8f7c04170544f5c4bef7a9af70ce0d7a25bfc4b4b8b1ba3e552413273bb7d"
|
| 841 |
+
},
|
| 842 |
+
"evidence": [
|
| 843 |
+
{
|
| 844 |
+
"path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/COMPLETE.json",
|
| 845 |
+
"sha256": "d7745d5ec1e08100629c55e2c7518ff408462e6adea4ebfb01b7b401b212c4b6"
|
| 846 |
+
},
|
| 847 |
+
{
|
| 848 |
+
"path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/METADATA.json",
|
| 849 |
+
"sha256": "6a522387522b92a0c01cbd8f32cdaa154289318a98564297d2f5723bc4c3dee0"
|
| 850 |
+
},
|
| 851 |
+
{
|
| 852 |
+
"path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/predictions.jsonl",
|
| 853 |
+
"sha256": "61bea4eb25014f6277cf81010872c1724e6d8f07a89803063649a21160b02edc"
|
| 854 |
+
},
|
| 855 |
+
{
|
| 856 |
+
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/REPORT.json",
|
| 857 |
+
"sha256": "1097d4eeb4abc792cac569e4793c469a332dc24bbdf5568a242ce4cee4c03759"
|
| 858 |
+
},
|
| 859 |
+
{
|
| 860 |
+
"path": "eval/open-reference-comparison-v1/REGISTRY.json",
|
| 861 |
+
"sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
|
| 862 |
+
},
|
| 863 |
+
{
|
| 864 |
+
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/COMPOSABLE-MANIFEST.json",
|
| 865 |
+
"sha256": "82a7589e866610251b5b52f5adb20ab5c7ed9fa6fb2a184c9d6a3367576171e7"
|
| 866 |
+
}
|
| 867 |
+
]
|
| 868 |
+
},
|
| 869 |
+
"nimble-9b": {
|
| 870 |
+
"panels": {
|
| 871 |
+
"old_core": {
|
| 872 |
+
"path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/nimble-9b/quality/normalized.jsonl",
|
| 873 |
+
"sha256": "6fefc71c6fbea1106894c1081057c2c7a253fa4ebc2e6aecfe23abc347be08bb"
|
| 874 |
+
},
|
| 875 |
+
"v3_core": {
|
| 876 |
+
"path": "eval/heldout/v3/baselines/nimble-9b/core/normalized.jsonl",
|
| 877 |
+
"sha256": "ab5323762fc5d34961c0fbce78497627528b696c60e18b880d671b468db97bf9"
|
| 878 |
+
},
|
| 879 |
+
"v4": {
|
| 880 |
+
"path": "results/open-reference-gap-v1/nimble-9b/v4/normalized.jsonl",
|
| 881 |
+
"sha256": "6ae263cef9cafdbcd35cd1f1dcc5b1bd920d3e6023d4fc487970286c04ec8d05"
|
| 882 |
+
},
|
| 883 |
+
"v5": {
|
| 884 |
+
"path": "eval/heldout/v5/open-baselines-v1/nimble-9b/core/normalized.jsonl",
|
| 885 |
+
"sha256": "0fc6dae372265ef62ea87ce9dc07b63a77ccd33e3aa48e15754f5ab98c32580a"
|
| 886 |
+
}
|
| 887 |
+
},
|
| 888 |
+
"transfer_rows": {
|
| 889 |
+
"path": "analysis/decision-benchmark-v3-baselines/nimble-9b/ROWS.json",
|
| 890 |
+
"sha256": "f55f29f973b0974adc7fc3fd582e2e5c0414f440e4a094a7ec8e152685e62a28"
|
| 891 |
+
},
|
| 892 |
+
"evidence": [
|
| 893 |
+
{
|
| 894 |
+
"path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/COMPLETE.json",
|
| 895 |
+
"sha256": "efc5a1be536227db99ff9937b84e9a5bc688fc497c0e4a84d8261144cea8a51d"
|
| 896 |
+
},
|
| 897 |
+
{
|
| 898 |
+
"path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/METADATA.json",
|
| 899 |
+
"sha256": "77eaa373ec6a74b483ac9b0cda739ecd8cce95a046266cd7a28488a768b82940"
|
| 900 |
+
},
|
| 901 |
+
{
|
| 902 |
+
"path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/predictions.jsonl",
|
| 903 |
+
"sha256": "ad2212fa691c662c7917b33315439844339252778fdaefcc3080cdf46e00a434"
|
| 904 |
+
},
|
| 905 |
+
{
|
| 906 |
+
"path": "analysis/decision-benchmark-v3-baselines/nimble-9b/REPORT.json",
|
| 907 |
+
"sha256": "b6065189b7522e5d0d2b7b088b673ba26212bee5fda2f0d7c6cd67fc9cc93136"
|
| 908 |
+
},
|
| 909 |
+
{
|
| 910 |
+
"path": "eval/open-reference-comparison-v1/REGISTRY.json",
|
| 911 |
+
"sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
|
| 912 |
+
},
|
| 913 |
+
{
|
| 914 |
+
"path": "analysis/decision-benchmark-v3-baselines/nimble-9b/COMPOSABLE-MANIFEST.json",
|
| 915 |
+
"sha256": "7423a2430984cf8e7eed2c5933ceeb0e2894d252e30e447d1294b6b6b08192bd"
|
| 916 |
+
}
|
| 917 |
+
]
|
| 918 |
+
}
|
| 919 |
+
},
|
| 920 |
+
"task_rows": 54,
|
| 921 |
+
"rank_public_models": [
|
| 922 |
+
"Jev",
|
| 923 |
+
"Lux",
|
| 924 |
+
"Nox",
|
| 925 |
+
"kev-9b",
|
| 926 |
+
"kev-4b",
|
| 927 |
+
"Qwen3.5-9B",
|
| 928 |
+
"Decider",
|
| 929 |
+
"Qwen3.5-4B",
|
| 930 |
+
"Sol",
|
| 931 |
+
"kev-0.8b",
|
| 932 |
+
"Qwen3.5-2B",
|
| 933 |
+
"Laya-base",
|
| 934 |
+
"Laya-multilingual",
|
| 935 |
+
"Kai"
|
| 936 |
+
],
|
| 937 |
+
"internal_extra_models_not_public_rank": [
|
| 938 |
+
"llm2jev-2b",
|
| 939 |
+
"llm2jev-4b",
|
| 940 |
+
"nimble-9b"
|
| 941 |
+
],
|
| 942 |
+
"laya_parameter_audit": {
|
| 943 |
+
"utc": "2026-09-22T05:01:58.295720+00:00",
|
| 944 |
+
"torch": "2.12.0+git6bbd260",
|
| 945 |
+
"transformers": "5.17.0",
|
| 946 |
+
"device": "meta, CPU-only process; no GPU devices exposed",
|
| 947 |
+
"models": {
|
| 948 |
+
"laya-english": {
|
| 949 |
+
"serialized_scalar_count": 421293830,
|
| 950 |
+
"unique_parameter_count": 421293827,
|
| 951 |
+
"parameter_entries": 205,
|
| 952 |
+
"named_parameter_entries_with_duplicates": 205,
|
| 953 |
+
"tied_parameter_aliases": [],
|
| 954 |
+
"persistent_buffers": {
|
| 955 |
+
"temperature": [
|
| 956 |
+
3
|
| 957 |
+
]
|
| 958 |
+
},
|
| 959 |
+
"persistent_buffer_scalar_count": 3,
|
| 960 |
+
"nonpersistent_buffers": {
|
| 961 |
+
"encoder.rotary_emb.full_attention_inv_freq": [
|
| 962 |
+
32
|
| 963 |
+
],
|
| 964 |
+
"encoder.rotary_emb.full_attention_original_inv_freq": [
|
| 965 |
+
32
|
| 966 |
+
],
|
| 967 |
+
"encoder.rotary_emb.sliding_attention_inv_freq": [
|
| 968 |
+
32
|
| 969 |
+
],
|
| 970 |
+
"encoder.rotary_emb.sliding_attention_original_inv_freq": [
|
| 971 |
+
32
|
| 972 |
+
]
|
| 973 |
+
},
|
| 974 |
+
"exact_state_dict_shapes_match_header": true,
|
| 975 |
+
"header_sha256": "3a42d8e7a96d5aa22c5bc9475d7ff092e32a832ae7c6973688f66f7a84048c04",
|
| 976 |
+
"source_sha256": "8d83611d480c971d640a7b7d3aa2f2219c5e8455e9cc2329fd073681bd8be23e",
|
| 977 |
+
"encoder_config_sha256": "bf3ab80598fdccf414855a2ce80f22859e4492d06ca8a62ddd1cfb63972f8979",
|
| 978 |
+
"agent_config_sha256": "ae287b56bbcf5f8c4f4541ae9dfd00c914c4c48b940b8398c3058af37ba92bbd",
|
| 979 |
+
"architecture": "ModernBertModel",
|
| 980 |
+
"interpretation": "Complete decision model, including encoder and heads; temperature calibration buffer excluded. No tied parameter aliases in the instantiated published model."
|
| 981 |
+
},
|
| 982 |
+
"laya-multilingual": {
|
| 983 |
+
"serialized_scalar_count": 321908998,
|
| 984 |
+
"unique_parameter_count": 321908995,
|
| 985 |
+
"parameter_entries": 169,
|
| 986 |
+
"named_parameter_entries_with_duplicates": 169,
|
| 987 |
+
"tied_parameter_aliases": [],
|
| 988 |
+
"persistent_buffers": {
|
| 989 |
+
"temperature": [
|
| 990 |
+
3
|
| 991 |
+
]
|
| 992 |
+
},
|
| 993 |
+
"persistent_buffer_scalar_count": 3,
|
| 994 |
+
"nonpersistent_buffers": {
|
| 995 |
+
"encoder.rotary_emb.full_attention_inv_freq": [
|
| 996 |
+
32
|
| 997 |
+
],
|
| 998 |
+
"encoder.rotary_emb.full_attention_original_inv_freq": [
|
| 999 |
+
32
|
| 1000 |
+
],
|
| 1001 |
+
"encoder.rotary_emb.sliding_attention_inv_freq": [
|
| 1002 |
+
32
|
| 1003 |
+
],
|
| 1004 |
+
"encoder.rotary_emb.sliding_attention_original_inv_freq": [
|
| 1005 |
+
32
|
| 1006 |
+
]
|
| 1007 |
+
},
|
| 1008 |
+
"exact_state_dict_shapes_match_header": true,
|
| 1009 |
+
"header_sha256": "22ab94329063133fdd2944997b906bb6076d2cfd3e43eb05d7ea92187e9c3984",
|
| 1010 |
+
"source_sha256": "8d83611d480c971d640a7b7d3aa2f2219c5e8455e9cc2329fd073681bd8be23e",
|
| 1011 |
+
"encoder_config_sha256": "83f6916d13ef0f556ac461f28308dc2bffa7ebeadee8ec9e2db5812020ea5bb4",
|
| 1012 |
+
"agent_config_sha256": "25061739243b617ad88d1219ba6f8a9c86c5881ca28df024fa2d9b3b2fcc30c6",
|
| 1013 |
+
"architecture": "ModernBertModel",
|
| 1014 |
+
"interpretation": "Complete decision model, including encoder and heads; temperature calibration buffer excluded. No tied parameter aliases in the instantiated published model."
|
| 1015 |
+
}
|
| 1016 |
+
},
|
| 1017 |
+
"source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
|
| 1018 |
+
}
|
| 1019 |
+
}
|
metrics/question-scaling.json
CHANGED
|
@@ -1,225 +1,16 @@
|
|
| 1 |
{
|
| 2 |
-
"
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
"measured_requests_per_second": 30.5029057383203,
|
| 11 |
-
"measured_questions_per_second": 30.5029057383203,
|
| 12 |
-
"max_allocated_bytes": 8649496576,
|
| 13 |
-
"max_reserved_bytes": 8923381760,
|
| 14 |
-
"input_tokens": 309
|
| 15 |
-
},
|
| 16 |
-
"choice-q2": {
|
| 17 |
-
"samples": 30,
|
| 18 |
-
"p50_ms": 33.3187065,
|
| 19 |
-
"p95_ms": 33.56022575,
|
| 20 |
-
"requests_per_second_from_median": 30.01316992903071,
|
| 21 |
-
"questions_per_second_from_median": 60.02633985806142,
|
| 22 |
-
"measured_requests_per_second": 30.051085192016746,
|
| 23 |
-
"measured_questions_per_second": 60.10217038403349,
|
| 24 |
-
"max_allocated_bytes": 8694344704,
|
| 25 |
-
"max_reserved_bytes": 8938061824,
|
| 26 |
-
"input_tokens": 618
|
| 27 |
-
},
|
| 28 |
-
"choice-q4": {
|
| 29 |
-
"samples": 30,
|
| 30 |
-
"p50_ms": 41.779055,
|
| 31 |
-
"p95_ms": 42.214628999999995,
|
| 32 |
-
"requests_per_second_from_median": 23.935438463124644,
|
| 33 |
-
"questions_per_second_from_median": 95.74175385249858,
|
| 34 |
-
"measured_requests_per_second": 23.884399734574895,
|
| 35 |
-
"measured_questions_per_second": 95.53759893829958,
|
| 36 |
-
"max_allocated_bytes": 8783385600,
|
| 37 |
-
"max_reserved_bytes": 9141485568,
|
| 38 |
-
"input_tokens": 1236
|
| 39 |
-
},
|
| 40 |
-
"choice-q8": {
|
| 41 |
-
"samples": 30,
|
| 42 |
-
"p50_ms": 69.641899,
|
| 43 |
-
"p95_ms": 70.72795845,
|
| 44 |
-
"requests_per_second_from_median": 14.359171911725154,
|
| 45 |
-
"questions_per_second_from_median": 114.87337529380123,
|
| 46 |
-
"measured_requests_per_second": 14.311513624064359,
|
| 47 |
-
"measured_questions_per_second": 114.49210899251487,
|
| 48 |
-
"max_allocated_bytes": 8963302400,
|
| 49 |
-
"max_reserved_bytes": 9376366592,
|
| 50 |
-
"input_tokens": 2472
|
| 51 |
-
},
|
| 52 |
-
"choice-q16": {
|
| 53 |
-
"samples": 30,
|
| 54 |
-
"p50_ms": 141.0197655,
|
| 55 |
-
"p95_ms": 141.85338695,
|
| 56 |
-
"requests_per_second_from_median": 7.091204530474134,
|
| 57 |
-
"questions_per_second_from_median": 113.45927248758615,
|
| 58 |
-
"measured_requests_per_second": 7.1075954042236935,
|
| 59 |
-
"measured_questions_per_second": 113.7215264675791,
|
| 60 |
-
"max_allocated_bytes": 8962254336,
|
| 61 |
-
"max_reserved_bytes": 9244246016,
|
| 62 |
-
"input_tokens": 4944
|
| 63 |
-
},
|
| 64 |
-
"choice-q32": {
|
| 65 |
-
"samples": 30,
|
| 66 |
-
"p50_ms": 282.0275835,
|
| 67 |
-
"p95_ms": 283.68355145,
|
| 68 |
-
"requests_per_second_from_median": 3.545752467151852,
|
| 69 |
-
"questions_per_second_from_median": 113.46407894885927,
|
| 70 |
-
"measured_requests_per_second": 3.5466119610851,
|
| 71 |
-
"measured_questions_per_second": 113.4915827547232,
|
| 72 |
-
"max_allocated_bytes": 8963302912,
|
| 73 |
-
"max_reserved_bytes": 9141485568,
|
| 74 |
-
"input_tokens": 9888
|
| 75 |
-
},
|
| 76 |
-
"noul-q1": {
|
| 77 |
-
"samples": 30,
|
| 78 |
-
"p50_ms": 32.685345,
|
| 79 |
-
"p95_ms": 33.000833899999996,
|
| 80 |
-
"requests_per_second_from_median": 30.594751256258732,
|
| 81 |
-
"questions_per_second_from_median": 30.594751256258732,
|
| 82 |
-
"measured_requests_per_second": 30.61657989325799,
|
| 83 |
-
"measured_questions_per_second": 30.61657989325799,
|
| 84 |
-
"max_allocated_bytes": 8645022720,
|
| 85 |
-
"max_reserved_bytes": 8917090304,
|
| 86 |
-
"input_tokens": 262
|
| 87 |
-
},
|
| 88 |
-
"noul-q2": {
|
| 89 |
-
"samples": 30,
|
| 90 |
-
"p50_ms": 32.822731000000005,
|
| 91 |
-
"p95_ms": 33.580505,
|
| 92 |
-
"requests_per_second_from_median": 30.466690903934833,
|
| 93 |
-
"questions_per_second_from_median": 60.933381807869665,
|
| 94 |
-
"measured_requests_per_second": 30.306271635312687,
|
| 95 |
-
"measured_questions_per_second": 60.61254327062537,
|
| 96 |
-
"max_allocated_bytes": 8686444544,
|
| 97 |
-
"max_reserved_bytes": 9216983040,
|
| 98 |
-
"input_tokens": 524
|
| 99 |
-
},
|
| 100 |
-
"noul-q4": {
|
| 101 |
-
"samples": 30,
|
| 102 |
-
"p50_ms": 38.8360825,
|
| 103 |
-
"p95_ms": 39.62854105,
|
| 104 |
-
"requests_per_second_from_median": 25.749250069185013,
|
| 105 |
-
"questions_per_second_from_median": 102.99700027674005,
|
| 106 |
-
"measured_requests_per_second": 25.627138601267134,
|
| 107 |
-
"measured_questions_per_second": 102.50855440506854,
|
| 108 |
-
"max_allocated_bytes": 8767323136,
|
| 109 |
-
"max_reserved_bytes": 9191817216,
|
| 110 |
-
"input_tokens": 1048
|
| 111 |
-
},
|
| 112 |
-
"noul-q8": {
|
| 113 |
-
"samples": 30,
|
| 114 |
-
"p50_ms": 64.278516,
|
| 115 |
-
"p95_ms": 65.07748314999999,
|
| 116 |
-
"requests_per_second_from_median": 15.557297558020787,
|
| 117 |
-
"questions_per_second_from_median": 124.4583804641663,
|
| 118 |
-
"measured_requests_per_second": 15.523243079930335,
|
| 119 |
-
"measured_questions_per_second": 124.18594463944268,
|
| 120 |
-
"max_allocated_bytes": 8932226048,
|
| 121 |
-
"max_reserved_bytes": 9141485568,
|
| 122 |
-
"input_tokens": 2096
|
| 123 |
-
},
|
| 124 |
-
"noul-q16": {
|
| 125 |
-
"samples": 30,
|
| 126 |
-
"p50_ms": 129.2654275,
|
| 127 |
-
"p95_ms": 130.59987905,
|
| 128 |
-
"requests_per_second_from_median": 7.7360205225794045,
|
| 129 |
-
"questions_per_second_from_median": 123.77632836127047,
|
| 130 |
-
"measured_requests_per_second": 7.725552951411574,
|
| 131 |
-
"measured_questions_per_second": 123.60884722258518,
|
| 132 |
-
"max_allocated_bytes": 8931440128,
|
| 133 |
-
"max_reserved_bytes": 9244246016,
|
| 134 |
-
"input_tokens": 4192
|
| 135 |
-
},
|
| 136 |
-
"noul-q32": {
|
| 137 |
-
"samples": 30,
|
| 138 |
-
"p50_ms": 259.156278,
|
| 139 |
-
"p95_ms": 261.59688335,
|
| 140 |
-
"requests_per_second_from_median": 3.8586755748977075,
|
| 141 |
-
"questions_per_second_from_median": 123.47761839672664,
|
| 142 |
-
"measured_requests_per_second": 3.8613862674541757,
|
| 143 |
-
"measured_questions_per_second": 123.56436055853362,
|
| 144 |
-
"max_allocated_bytes": 8932226560,
|
| 145 |
-
"max_reserved_bytes": 9244246016,
|
| 146 |
-
"input_tokens": 8384
|
| 147 |
-
},
|
| 148 |
-
"score-q1": {
|
| 149 |
-
"samples": 30,
|
| 150 |
-
"p50_ms": 33.319307,
|
| 151 |
-
"p95_ms": 34.269436,
|
| 152 |
-
"requests_per_second_from_median": 30.01262901416287,
|
| 153 |
-
"questions_per_second_from_median": 30.01262901416287,
|
| 154 |
-
"measured_requests_per_second": 29.995543652066246,
|
| 155 |
-
"measured_questions_per_second": 29.995543652066246,
|
| 156 |
-
"max_allocated_bytes": 8648972288,
|
| 157 |
-
"max_reserved_bytes": 9244246016,
|
| 158 |
-
"input_tokens": 307
|
| 159 |
-
},
|
| 160 |
-
"score-q2": {
|
| 161 |
-
"samples": 30,
|
| 162 |
-
"p50_ms": 33.3555535,
|
| 163 |
-
"p95_ms": 33.76658145,
|
| 164 |
-
"requests_per_second_from_median": 29.980015171986278,
|
| 165 |
-
"questions_per_second_from_median": 59.960030343972555,
|
| 166 |
-
"measured_requests_per_second": 30.07546038274141,
|
| 167 |
-
"measured_questions_per_second": 60.15092076548282,
|
| 168 |
-
"max_allocated_bytes": 8693689344,
|
| 169 |
-
"max_reserved_bytes": 9141485568,
|
| 170 |
-
"input_tokens": 614
|
| 171 |
-
},
|
| 172 |
-
"score-q4": {
|
| 173 |
-
"samples": 30,
|
| 174 |
-
"p50_ms": 41.8594245,
|
| 175 |
-
"p95_ms": 42.2104965,
|
| 176 |
-
"requests_per_second_from_median": 23.889482761522437,
|
| 177 |
-
"questions_per_second_from_median": 95.55793104608975,
|
| 178 |
-
"measured_requests_per_second": 23.829400385190663,
|
| 179 |
-
"measured_questions_per_second": 95.31760154076265,
|
| 180 |
-
"max_allocated_bytes": 8782337024,
|
| 181 |
-
"max_reserved_bytes": 9244246016,
|
| 182 |
-
"input_tokens": 1228
|
| 183 |
-
},
|
| 184 |
-
"score-q8": {
|
| 185 |
-
"samples": 30,
|
| 186 |
-
"p50_ms": 69.734315,
|
| 187 |
-
"p95_ms": 70.12355445,
|
| 188 |
-
"requests_per_second_from_median": 14.340142295797987,
|
| 189 |
-
"questions_per_second_from_median": 114.7211383663839,
|
| 190 |
-
"measured_requests_per_second": 14.393887061255466,
|
| 191 |
-
"measured_questions_per_second": 115.15109649004373,
|
| 192 |
-
"max_allocated_bytes": 8963302400,
|
| 193 |
-
"max_reserved_bytes": 9376366592,
|
| 194 |
-
"input_tokens": 2456
|
| 195 |
-
},
|
| 196 |
-
"score-q16": {
|
| 197 |
-
"samples": 30,
|
| 198 |
-
"p50_ms": 141.064813,
|
| 199 |
-
"p95_ms": 141.96145435,
|
| 200 |
-
"requests_per_second_from_median": 7.088940032125517,
|
| 201 |
-
"questions_per_second_from_median": 113.42304051400828,
|
| 202 |
-
"measured_requests_per_second": 7.0981910052955834,
|
| 203 |
-
"measured_questions_per_second": 113.57105608472934,
|
| 204 |
-
"max_allocated_bytes": 8963302912,
|
| 205 |
-
"max_reserved_bytes": 9244246016,
|
| 206 |
-
"input_tokens": 4912
|
| 207 |
-
},
|
| 208 |
-
"score-q32": {
|
| 209 |
-
"samples": 30,
|
| 210 |
-
"p50_ms": 282.3302425,
|
| 211 |
-
"p95_ms": 285.68878665,
|
| 212 |
-
"requests_per_second_from_median": 3.541951408198858,
|
| 213 |
-
"questions_per_second_from_median": 113.34244506236345,
|
| 214 |
-
"measured_requests_per_second": 3.5352975301699447,
|
| 215 |
-
"measured_questions_per_second": 113.12952096543823,
|
| 216 |
-
"max_allocated_bytes": 8963302912,
|
| 217 |
-
"max_reserved_bytes": 9244246016,
|
| 218 |
-
"input_tokens": 9824
|
| 219 |
-
}
|
| 220 |
-
}
|
| 221 |
},
|
| 222 |
-
"
|
|
|
|
|
|
|
| 223 |
1,
|
| 224 |
2,
|
| 225 |
4,
|
|
@@ -227,14 +18,96 @@
|
|
| 227 |
16,
|
| 228 |
32
|
| 229 |
],
|
| 230 |
-
"
|
| 231 |
-
|
| 232 |
-
|
| 233 |
-
|
| 234 |
-
|
| 235 |
-
|
| 236 |
-
|
| 237 |
-
|
| 238 |
-
|
| 239 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 240 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"model": "Nox",
|
| 3 |
+
"release": "v1.3.1",
|
| 4 |
+
"release_binding": {
|
| 5 |
+
"repo_id": "llm-semantic-router/Decision-1.0-Nox",
|
| 6 |
+
"revision": "74979f7c7408325716dcb89b686d9cf94301923a",
|
| 7 |
+
"release_tag": "v1.3.1",
|
| 8 |
+
"bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
|
| 9 |
+
"publication_receipt_sha256": "7d680d3a1dc5919fc2975cbea9ea5650f39cdbf317d8e1ca6b2d03d5ae11bcdb"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
},
|
| 11 |
+
"condition": "distinct_primary",
|
| 12 |
+
"type": "choice",
|
| 13 |
+
"Q": [
|
| 14 |
1,
|
| 15 |
2,
|
| 16 |
4,
|
|
|
|
| 18 |
16,
|
| 19 |
32
|
| 20 |
],
|
| 21 |
+
"tokens_per_question": 499,
|
| 22 |
+
"samples_per_point": 30,
|
| 23 |
+
"fresh_process_blocks": 6,
|
| 24 |
+
"points": {
|
| 25 |
+
"1": {
|
| 26 |
+
"count": 30,
|
| 27 |
+
"p50_ms": 32.678454,
|
| 28 |
+
"p95_ms": 33.71943,
|
| 29 |
+
"mean_ms": 32.7633463,
|
| 30 |
+
"std_ms": 0.5253808729894747,
|
| 31 |
+
"minimum_ms": 31.98487,
|
| 32 |
+
"maximum_ms": 33.745036,
|
| 33 |
+
"questions_per_second_at_p50": 30.601202859841532,
|
| 34 |
+
"peak_allocated_bytes": 8675011072,
|
| 35 |
+
"peak_reserved_bytes": 9246343168,
|
| 36 |
+
"questions": 1,
|
| 37 |
+
"tokens_per_question": 499
|
| 38 |
+
},
|
| 39 |
+
"2": {
|
| 40 |
+
"count": 30,
|
| 41 |
+
"p50_ms": 36.370171,
|
| 42 |
+
"p95_ms": 36.64674145,
|
| 43 |
+
"mean_ms": 36.389989299999996,
|
| 44 |
+
"std_ms": 0.11128765575847639,
|
| 45 |
+
"minimum_ms": 36.268881,
|
| 46 |
+
"maximum_ms": 36.692332,
|
| 47 |
+
"questions_per_second_at_p50": 54.99011813829525,
|
| 48 |
+
"peak_allocated_bytes": 8746551808,
|
| 49 |
+
"peak_reserved_bytes": 9246343168,
|
| 50 |
+
"questions": 2,
|
| 51 |
+
"tokens_per_question": 499
|
| 52 |
+
},
|
| 53 |
+
"4": {
|
| 54 |
+
"count": 30,
|
| 55 |
+
"p50_ms": 60.415027,
|
| 56 |
+
"p95_ms": 60.9524315,
|
| 57 |
+
"mean_ms": 60.47764396666667,
|
| 58 |
+
"std_ms": 0.286383903775007,
|
| 59 |
+
"minimum_ms": 60.134886,
|
| 60 |
+
"maximum_ms": 61.481055,
|
| 61 |
+
"questions_per_second_at_p50": 66.20869341000211,
|
| 62 |
+
"peak_allocated_bytes": 8890681856,
|
| 63 |
+
"peak_reserved_bytes": 9246343168,
|
| 64 |
+
"questions": 4,
|
| 65 |
+
"tokens_per_question": 499
|
| 66 |
+
},
|
| 67 |
+
"8": {
|
| 68 |
+
"count": 30,
|
| 69 |
+
"p50_ms": 103.619237,
|
| 70 |
+
"p95_ms": 104.1805109,
|
| 71 |
+
"mean_ms": 103.640764,
|
| 72 |
+
"std_ms": 0.34493524425007943,
|
| 73 |
+
"minimum_ms": 102.972861,
|
| 74 |
+
"maximum_ms": 104.235019,
|
| 75 |
+
"questions_per_second_at_p50": 77.20574124667604,
|
| 76 |
+
"peak_allocated_bytes": 9178941952,
|
| 77 |
+
"peak_reserved_bytes": 9246343168,
|
| 78 |
+
"questions": 8,
|
| 79 |
+
"tokens_per_question": 499
|
| 80 |
+
},
|
| 81 |
+
"16": {
|
| 82 |
+
"count": 30,
|
| 83 |
+
"p50_ms": 206.9910775,
|
| 84 |
+
"p95_ms": 207.62369065000001,
|
| 85 |
+
"mean_ms": 206.9661473,
|
| 86 |
+
"std_ms": 0.49304760661506103,
|
| 87 |
+
"minimum_ms": 206.100008,
|
| 88 |
+
"maximum_ms": 207.740005,
|
| 89 |
+
"questions_per_second_at_p50": 77.2980178336431,
|
| 90 |
+
"peak_allocated_bytes": 9178946560,
|
| 91 |
+
"peak_reserved_bytes": 9246343168,
|
| 92 |
+
"questions": 16,
|
| 93 |
+
"tokens_per_question": 499
|
| 94 |
+
},
|
| 95 |
+
"32": {
|
| 96 |
+
"count": 30,
|
| 97 |
+
"p50_ms": 414.19944399999997,
|
| 98 |
+
"p95_ms": 415.3707452,
|
| 99 |
+
"mean_ms": 414.2741519666667,
|
| 100 |
+
"std_ms": 0.6025438072196239,
|
| 101 |
+
"minimum_ms": 413.162275,
|
| 102 |
+
"maximum_ms": 416.066377,
|
| 103 |
+
"questions_per_second_at_p50": 77.25746729877311,
|
| 104 |
+
"peak_allocated_bytes": 9178946560,
|
| 105 |
+
"peak_reserved_bytes": 9246343168,
|
| 106 |
+
"questions": 32,
|
| 107 |
+
"tokens_per_question": 499
|
| 108 |
+
}
|
| 109 |
+
},
|
| 110 |
+
"timing_summary_sha256": "501f25efd1813b566d77e98278229440b7a3b4a2358def1462c0171a33e9f3be",
|
| 111 |
+
"API_sha256": "273f6f10f22d5a68b8db34cfcbd35407fb43d8030f6d7cd188bcf118dd90a152",
|
| 112 |
+
"scope": "Only optimized path subsequently published, no old-release series. Fixed-length Python requests, excludes network and loading."
|
| 113 |
}
|
release-manifest.json
CHANGED
|
@@ -1,12 +1,12 @@
|
|
| 1 |
{
|
| 2 |
"format": "decision-public-release-v1",
|
| 3 |
-
"status": "
|
| 4 |
"bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
|
| 5 |
-
"readiness_sha256": "
|
| 6 |
-
"model_card_sha256": "
|
| 7 |
"repo_id": "llm-semantic-router/Decision-1.0-Nox",
|
| 8 |
-
"assembly_script_sha256": "
|
| 9 |
-
"original_bundle_manifest_preserved":
|
| 10 |
"files_exclude_this_manifest": true,
|
| 11 |
"files": [
|
| 12 |
{
|
|
@@ -14,6 +14,11 @@
|
|
| 14 |
"bytes": 6606,
|
| 15 |
"sha256": "e9ed3c41423a6d15edbf82e14b80897fe5ef6cd454bcd9957d91b31460461718"
|
| 16 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
{
|
| 18 |
"file": "Dockerfile.runtime",
|
| 19 |
"bytes": 751,
|
|
@@ -21,8 +26,8 @@
|
|
| 21 |
},
|
| 22 |
{
|
| 23 |
"file": "EVALUATION.md",
|
| 24 |
-
"bytes":
|
| 25 |
-
"sha256": "
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"file": "LICENSE",
|
|
@@ -31,8 +36,8 @@
|
|
| 31 |
},
|
| 32 |
{
|
| 33 |
"file": "MATERIALS.json",
|
| 34 |
-
"bytes":
|
| 35 |
-
"sha256": "
|
| 36 |
},
|
| 37 |
{
|
| 38 |
"file": "NORMALIZATION_RUNTIME.md",
|
|
@@ -41,8 +46,8 @@
|
|
| 41 |
},
|
| 42 |
{
|
| 43 |
"file": "QUESTION-SCALING.md",
|
| 44 |
-
"bytes":
|
| 45 |
-
"sha256": "
|
| 46 |
},
|
| 47 |
{
|
| 48 |
"file": "QWEN-LICENSE",
|
|
@@ -51,8 +56,8 @@
|
|
| 51 |
},
|
| 52 |
{
|
| 53 |
"file": "README.md",
|
| 54 |
-
"bytes":
|
| 55 |
-
"sha256": "
|
| 56 |
},
|
| 57 |
{
|
| 58 |
"file": "RUNTIME-RELEASE.json",
|
|
@@ -69,6 +74,11 @@
|
|
| 69 |
"bytes": 2621,
|
| 70 |
"sha256": "fb76d9fc9147e15a678eb91dfc40d7039845a4eaebdecce041a5c9e76c3d3e0f"
|
| 71 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 72 |
{
|
| 73 |
"file": "SERVING_OPTIMIZATION.json",
|
| 74 |
"bytes": 1548,
|
|
@@ -79,6 +89,11 @@
|
|
| 79 |
"bytes": 5637,
|
| 80 |
"sha256": "d437bc0149bcb8c9891fbb33c5abc4c2336c3a89ab6cfa0981da5a7f1d19f1f4"
|
| 81 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
{
|
| 83 |
"file": "USAGE.md",
|
| 84 |
"bytes": 5435,
|
|
@@ -219,6 +234,21 @@
|
|
| 219 |
"bytes": 2962868,
|
| 220 |
"sha256": "213511289ce8df038d938ac470e803c427ed57f0f85cc397dd4d79964b866541"
|
| 221 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 222 |
{
|
| 223 |
"file": "assets/decision-question-scaling-600px.png",
|
| 224 |
"bytes": 40825,
|
|
@@ -226,18 +256,33 @@
|
|
| 226 |
},
|
| 227 |
{
|
| 228 |
"file": "assets/decision-question-scaling.pdf",
|
| 229 |
-
"bytes":
|
| 230 |
-
"sha256": "
|
| 231 |
},
|
| 232 |
{
|
| 233 |
"file": "assets/decision-question-scaling.png",
|
| 234 |
-
"bytes":
|
| 235 |
-
"sha256": "
|
| 236 |
},
|
| 237 |
{
|
| 238 |
"file": "assets/decision-question-scaling.svg",
|
| 239 |
-
"bytes":
|
| 240 |
-
"sha256": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 241 |
},
|
| 242 |
{
|
| 243 |
"file": "assets/readout.png",
|
|
@@ -314,11 +359,21 @@
|
|
| 314 |
"bytes": 10529624,
|
| 315 |
"sha256": "9cb6f639714e31bcb76b58eaf94af0b72ac3db9d091ebcda45d9b575efd489de"
|
| 316 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 317 |
{
|
| 318 |
"file": "metrics/comparator-coverage.json",
|
| 319 |
"bytes": 24628,
|
| 320 |
"sha256": "fb4e6c8deee00798056aff95ceaaec4f074effbd8b84c20e69667869b1a95c03"
|
| 321 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 322 |
{
|
| 323 |
"file": "metrics/expanded-quality.json",
|
| 324 |
"bytes": 64295,
|
|
@@ -336,8 +391,8 @@
|
|
| 336 |
},
|
| 337 |
{
|
| 338 |
"file": "metrics/question-scaling.json",
|
| 339 |
-
"bytes":
|
| 340 |
-
"sha256": "
|
| 341 |
},
|
| 342 |
{
|
| 343 |
"file": "metrics/semantic-consistency.json",
|
|
@@ -502,9 +557,9 @@
|
|
| 502 |
"MATERIALS.json": "copy",
|
| 503 |
"metrics/semantic-consistency.json": "copy"
|
| 504 |
},
|
| 505 |
-
"scope": "
|
| 506 |
"release_tag": "v1.3.1",
|
| 507 |
-
"change_kind": "
|
| 508 |
-
"previous_main_revision": "
|
| 509 |
-
"previous_release_manifest_sha256": "
|
| 510 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"format": "decision-public-release-v1",
|
| 3 |
+
"status": "documentation-only-assembled",
|
| 4 |
"bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
|
| 5 |
+
"readiness_sha256": "fcbb84a9bf7294ac79e2421cdd2cbfdfc46a79fd74c41ba0965778472d1cafc1",
|
| 6 |
+
"model_card_sha256": "108a5727cda42f93dcc2a3d5fbe653cc42a5fccbe1574ad940553aafe2f1741d",
|
| 7 |
"repo_id": "llm-semantic-router/Decision-1.0-Nox",
|
| 8 |
+
"assembly_script_sha256": "1697c4abaee22bb7b60259c042488178fc71847afe4810037683f874d7b8d513",
|
| 9 |
+
"original_bundle_manifest_preserved": true,
|
| 10 |
"files_exclude_this_manifest": true,
|
| 11 |
"files": [
|
| 12 |
{
|
|
|
|
| 14 |
"bytes": 6606,
|
| 15 |
"sha256": "e9ed3c41423a6d15edbf82e14b80897fe5ef6cd454bcd9957d91b31460461718"
|
| 16 |
},
|
| 17 |
+
{
|
| 18 |
+
"file": "DIAGNOSTICS.md",
|
| 19 |
+
"bytes": 5967,
|
| 20 |
+
"sha256": "5492ed49869e45e45650c2501d02a78f9a2fbc17267c54f90cf034a01ebda2b3"
|
| 21 |
+
},
|
| 22 |
{
|
| 23 |
"file": "Dockerfile.runtime",
|
| 24 |
"bytes": 751,
|
|
|
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"file": "EVALUATION.md",
|
| 29 |
+
"bytes": 3744,
|
| 30 |
+
"sha256": "3866dea6ae78b09bf93ad7b843d067960493217ae84b6ded87bfdac97b424608"
|
| 31 |
},
|
| 32 |
{
|
| 33 |
"file": "LICENSE",
|
|
|
|
| 36 |
},
|
| 37 |
{
|
| 38 |
"file": "MATERIALS.json",
|
| 39 |
+
"bytes": 2688,
|
| 40 |
+
"sha256": "387ae91720084aefff2ce1964d30550485666587e7ea3a0a2bb5fae5af678d8d"
|
| 41 |
},
|
| 42 |
{
|
| 43 |
"file": "NORMALIZATION_RUNTIME.md",
|
|
|
|
| 46 |
},
|
| 47 |
{
|
| 48 |
"file": "QUESTION-SCALING.md",
|
| 49 |
+
"bytes": 1274,
|
| 50 |
+
"sha256": "898f3554434284ff506dbc63c34859750f07b33dcf71f5b6d4ec7a798a56eb28"
|
| 51 |
},
|
| 52 |
{
|
| 53 |
"file": "QWEN-LICENSE",
|
|
|
|
| 56 |
},
|
| 57 |
{
|
| 58 |
"file": "README.md",
|
| 59 |
+
"bytes": 4737,
|
| 60 |
+
"sha256": "108a5727cda42f93dcc2a3d5fbe653cc42a5fccbe1574ad940553aafe2f1741d"
|
| 61 |
},
|
| 62 |
{
|
| 63 |
"file": "RUNTIME-RELEASE.json",
|
|
|
|
| 74 |
"bytes": 2621,
|
| 75 |
"sha256": "fb76d9fc9147e15a678eb91dfc40d7039845a4eaebdecce041a5c9e76c3d3e0f"
|
| 76 |
},
|
| 77 |
+
{
|
| 78 |
+
"file": "SENSITIVITY.md",
|
| 79 |
+
"bytes": 783,
|
| 80 |
+
"sha256": "e03ec18140d5c18e606da23cc5b26ef41176ab38019912d85050677e4637f0a0"
|
| 81 |
+
},
|
| 82 |
{
|
| 83 |
"file": "SERVING_OPTIMIZATION.json",
|
| 84 |
"bytes": 1548,
|
|
|
|
| 89 |
"bytes": 5637,
|
| 90 |
"sha256": "d437bc0149bcb8c9891fbb33c5abc4c2336c3a89ab6cfa0981da5a7f1d19f1f4"
|
| 91 |
},
|
| 92 |
+
{
|
| 93 |
+
"file": "TASKS.md",
|
| 94 |
+
"bytes": 9784,
|
| 95 |
+
"sha256": "e87c5271049f09ce23fcb7fdcbaa9f03e5c50039b6d7a4005a6b4c7fbc1d30ca"
|
| 96 |
+
},
|
| 97 |
{
|
| 98 |
"file": "USAGE.md",
|
| 99 |
"bytes": 5435,
|
|
|
|
| 234 |
"bytes": 2962868,
|
| 235 |
"sha256": "213511289ce8df038d938ac470e803c427ed57f0f85cc397dd4d79964b866541"
|
| 236 |
},
|
| 237 |
+
{
|
| 238 |
+
"file": "assets/decision-matrix.pdf",
|
| 239 |
+
"bytes": 28634,
|
| 240 |
+
"sha256": "d1651d91db893ffbdf3261ec476241eb03c18c7a769d19978c668d60a1ac844f"
|
| 241 |
+
},
|
| 242 |
+
{
|
| 243 |
+
"file": "assets/decision-matrix.png",
|
| 244 |
+
"bytes": 345190,
|
| 245 |
+
"sha256": "b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc"
|
| 246 |
+
},
|
| 247 |
+
{
|
| 248 |
+
"file": "assets/decision-matrix.svg",
|
| 249 |
+
"bytes": 46970,
|
| 250 |
+
"sha256": "dc92d1e3f3cf0518ee44f1bc512aae74732e1db82b919774a52c0da7c744b0b4"
|
| 251 |
+
},
|
| 252 |
{
|
| 253 |
"file": "assets/decision-question-scaling-600px.png",
|
| 254 |
"bytes": 40825,
|
|
|
|
| 256 |
},
|
| 257 |
{
|
| 258 |
"file": "assets/decision-question-scaling.pdf",
|
| 259 |
+
"bytes": 18958,
|
| 260 |
+
"sha256": "d756f60c142ecaf647e59ea5bb3a611a39d1572a6471cd49afb7ef602a23fb3b"
|
| 261 |
},
|
| 262 |
{
|
| 263 |
"file": "assets/decision-question-scaling.png",
|
| 264 |
+
"bytes": 103662,
|
| 265 |
+
"sha256": "722e664974b28324693f5f4af6609330c278385ac118ed6f757b8fc0afec64ac"
|
| 266 |
},
|
| 267 |
{
|
| 268 |
"file": "assets/decision-question-scaling.svg",
|
| 269 |
+
"bytes": 9265,
|
| 270 |
+
"sha256": "cc591eafa335b5d16edd306d6368679afd283a4b14d5f83fc77780a352b1d315"
|
| 271 |
+
},
|
| 272 |
+
{
|
| 273 |
+
"file": "assets/decision-ranking.pdf",
|
| 274 |
+
"bytes": 24599,
|
| 275 |
+
"sha256": "f3e037ada6c30871d9ebe0087fa3d122346f7ab1c927515878d524c492b7799a"
|
| 276 |
+
},
|
| 277 |
+
{
|
| 278 |
+
"file": "assets/decision-ranking.png",
|
| 279 |
+
"bytes": 226418,
|
| 280 |
+
"sha256": "1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0"
|
| 281 |
+
},
|
| 282 |
+
{
|
| 283 |
+
"file": "assets/decision-ranking.svg",
|
| 284 |
+
"bytes": 14935,
|
| 285 |
+
"sha256": "c8de982ba51fb9d2b729191018220964426a0f524a00a796553b8879db271d1a"
|
| 286 |
},
|
| 287 |
{
|
| 288 |
"file": "assets/readout.png",
|
|
|
|
| 359 |
"bytes": 10529624,
|
| 360 |
"sha256": "9cb6f639714e31bcb76b58eaf94af0b72ac3db9d091ebcda45d9b575efd489de"
|
| 361 |
},
|
| 362 |
+
{
|
| 363 |
+
"file": "metrics/benchmark.json",
|
| 364 |
+
"bytes": 2451569,
|
| 365 |
+
"sha256": "50dac7884849cbdb8be58c1c82b811f51543a08a548253cb8295f4e6444c4718"
|
| 366 |
+
},
|
| 367 |
{
|
| 368 |
"file": "metrics/comparator-coverage.json",
|
| 369 |
"bytes": 24628,
|
| 370 |
"sha256": "fb4e6c8deee00798056aff95ceaaec4f074effbd8b84c20e69667869b1a95c03"
|
| 371 |
},
|
| 372 |
+
{
|
| 373 |
+
"file": "metrics/evaluation-provenance.json",
|
| 374 |
+
"bytes": 44403,
|
| 375 |
+
"sha256": "1f9c8b1e03d521160a8426798394418748daa72014dfc7272578d1e83fe237b4"
|
| 376 |
+
},
|
| 377 |
{
|
| 378 |
"file": "metrics/expanded-quality.json",
|
| 379 |
"bytes": 64295,
|
|
|
|
| 391 |
},
|
| 392 |
{
|
| 393 |
"file": "metrics/question-scaling.json",
|
| 394 |
+
"bytes": 3484,
|
| 395 |
+
"sha256": "9407f7ebf86581cdea21e704cdbebb9ced0a485eb980925f48233ebea0f830dc"
|
| 396 |
},
|
| 397 |
{
|
| 398 |
"file": "metrics/semantic-consistency.json",
|
|
|
|
| 557 |
"MATERIALS.json": "copy",
|
| 558 |
"metrics/semantic-consistency.json": "copy"
|
| 559 |
},
|
| 560 |
+
"scope": "Current14model five-panel product comparison,54task details, separate diagnostics,current API latency. Weights, runtime, tokenizer and calibration unchanged.",
|
| 561 |
"release_tag": "v1.3.1",
|
| 562 |
+
"change_kind": "documentation-only-current-benchmark",
|
| 563 |
+
"previous_main_revision": "74979f7c7408325716dcb89b686d9cf94301923a",
|
| 564 |
+
"previous_release_manifest_sha256": "7f54061edbe5e3801107a3792d07aaf94bf51e6e1b176b3e457c86b3477134ee"
|
| 565 |
}
|