Xunzhuo commited on
Commit
a3527b0
·
verified ·
1 Parent(s): 3e46eb1

Update public comparison roster; preserve model weights and benchmark scores

Browse files
DIAGNOSTICS.md CHANGED
@@ -19,7 +19,6 @@ These axes remain separate from headline accuracy. Probability metrics use the s
19
  | Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
20
  | Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
21
  | Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
22
- | Kai | 1046/1046 | 0.6066 | 1.1552 | 7.95 | 1.24 | 0.3343 |
23
 
24
  Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
25
 
@@ -40,7 +39,6 @@ Coverage at an error threshold keeps whole confidence-tie groups together. These
40
  | Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
41
  | Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
42
  | Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
43
- | Kai | 36/36 | 41.67 | 27.78 | 0.1170 |
44
 
45
  The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
46
 
@@ -61,7 +59,6 @@ The 36 paired permutations test the same semantics under changed option order. S
61
  | Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
62
  | Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
63
  | Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
64
- | Kai | 110/110 | 37.27 | 55.39 | 18.18 | 0.8262 | -0.07 |
65
 
66
  The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
67
 
@@ -82,9 +79,8 @@ The 110 unknowable examples have no scored true class and are excluded from accu
82
  | Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
83
  | Laya · English | 2720/2720 | 1264/1264 | 34 |
84
  | Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
85
- | Kai | 2720/2720 | 1264/1264 | 0 |
86
 
87
- Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation; Kai retains the shipped complete-request 1,024-token limit. Server-side truncation for Jev cannot be observed.
88
 
89
  ## Uncertainty
90
 
@@ -103,6 +99,5 @@ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and orde
103
  | Qwen3.5-2B | 57.24 | 55.74–58.76 |
104
  | Laya · English | 51.03 | 49.43–52.68 |
105
  | Laya · Multilingual | 47.19 | 45.58–48.82 |
106
- | Kai | 46.49 | 45.02–47.99 |
107
 
108
  Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
 
19
  | Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
20
  | Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
21
  | Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
 
22
 
23
  Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
24
 
 
39
  | Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
40
  | Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
41
  | Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
 
42
 
43
  The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
44
 
 
59
  | Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
60
  | Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
61
  | Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
 
62
 
63
  The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
64
 
 
79
  | Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
80
  | Laya · English | 2720/2720 | 1264/1264 | 34 |
81
  | Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
 
82
 
83
+ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation. Server-side truncation for Jev cannot be observed.
84
 
85
  ## Uncertainty
86
 
 
99
  | Qwen3.5-2B | 57.24 | 55.74–58.76 |
100
  | Laya · English | 51.03 | 49.43–52.68 |
101
  | Laya · Multilingual | 47.19 | 45.58–48.82 |
 
102
 
103
  Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
EVALUATION.md CHANGED
@@ -1,8 +1,6 @@
1
  # Evaluation
2
 
3
- Lux achieves **76.72%** on the decision-focused benchmark. The headline combines 880 general decisions, 880 compositional tasks, 480 natural-reading questions, 480 reading/inference questions and 1,046 answerable transfer decisions. The exact weights are 30%, 25%, 15%, 15% and 15%, preserving each original panel's family/source weights. Refusals and failures remain in the requested accuracy denominator.
4
-
5
- **These product weights were chosen after earlier results were observed.** Reweighting changes the product score, not model predictions or trained capability. [Weight sensitivity](WEIGHTING.md) retains both the earlier five-panel weighting and the original four-panel comparison for these same current model predictions.
6
 
7
  | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
8
  |---|---:|---:|---:|---:|---:|---:|---:|
@@ -19,24 +17,21 @@ Lux achieves **76.72%** on the decision-focused benchmark. The headline combines
19
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
20
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
21
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
22
- | Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
23
-
24
- Models are ordered by unrounded overall accuracy. Bold identifies a Decision-family result strictly higher than every external open reference in that column, excluding other Decision-family models and the closed Jev frontier. Ties are not bold. Sizes identify the source-model tier; Lux's deployed text-plus-head parameter count is 7.941B. Kai is a 0.572B encoder. Laya English and multilingual are evaluated separately.
25
 
26
- ## What the benchmark measures
27
 
28
- The original panels cover 27 tasks spanning routing, supplied-evidence decisions, state and rule composition, natural reading and inference. The transfer projection adds source generalization, controlled evidence transformations and answerable counterfactual controls from the audited Kev test suites. Unknown-information probes, option permutation and semantic consistency are separate diagnostics; they are not assigned fabricated correctness labels or silently mixed into the headline.
29
 
30
- [All 54 task results](TASKS.md) and [robustness diagnostics](DIAGNOSTICS.md) preserve the individual capabilities instead of hiding them behind the five aggregate columns. [Machine-readable benchmark and uncertainty](metrics/benchmark.json) includes the exact component scores, paired comparisons and comparator identities. Probability scores require complete valid distributions; unsupported rows do not receive invented probabilities.
31
 
32
- Confidence intervals use 10,000 paired bootstrap draws over source components, preserving the frozen within-panel weighting. They describe this finite benchmark and do not establish universal superiority. These suites are observed regression evidence; exclusion from custom training does not prove absence from upstream pretraining.
33
 
34
- ## Compared inference paths
35
 
36
- All models receive the same frozen request population through their recorded adapters. Native support limits are retained, including Kai's complete-input admission limit and reference-model option limits. Untuned Qwen models use a fixed pretrained vocabulary-head letter readout at temperature one, without decision-head training or generated reasoning. Lux uses its exported BF16 text backbone, FP32 candidate head and independent-CAL temperature. Jev is the recorded hosted service snapshot, not a locally inspectable architecture.
37
 
38
- Lux was selected and calibrated before this benchmark was inspected. No test result reselected another checkpoint. The portable public loader reproduced all 1,600 calibration raw-logit vectors exactly in an isolated container without original training-weight access. Exact source, model and runtime identities are retained in the bundle metadata.
39
 
40
- ## Measured latency
41
 
42
- [Question-count scaling](QUESTION-SCALING.md) measures Lux's current native local API with fixed input length per question. Its host differs from the earlier Sol/Nox timing host, so the curves are not combined into a cross-hardware speed ranking. Loading and network transport are excluded; tokenization, model execution and response construction are included.
 
1
  # Evaluation
2
 
3
+ The comparison covers **3,766 scored decisions across 54 tasks** and all 13 displayed models. All models use the same requested examples, candidate order and scoring rules. Each model keeps its native admission, truncation and probability semantics.
 
 
4
 
5
  | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
6
  |---|---:|---:|---:|---:|---:|---:|---:|
 
17
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
18
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
19
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
 
 
 
20
 
21
+ Accuracy (%). Rows are sorted by overall score. Bold marks Decision-family cells strictly above every external open or untuned reference in that metric; Jev and the other Decision models are excluded from that threshold.
22
 
23
+ ## Scope and weighting
24
 
25
+ The overall score weights **Decisions 30%, Composition 25%, Reading 15%, Inference 15%, Transfer 15%**. The first four panels contain 880, 880, 480 and 480 questions and retain their original family/source weights. Transfer is micro-accuracy over 1,046 clean knowable questions from the frozen upstream transfer test; 110 missing-evidence questions and 108 variants remain separate diagnostics. No latency, calibration error or consistency score is averaged into accuracy.
26
 
27
+ These are **outcome-informed product-priority weights, chosen after observing benchmark results**. The data are observed regression tests, not a fresh blind test. Reweighting is not a training improvement. [Weight sensitivity](SENSITIVITY.md) retains the prior weighting and original four-panel comparison for the same model weights. Training, checkpoint selection and calibration do not use these test labels.
28
 
29
+ ## Full results
30
 
31
+ [All 54 task rows](TASKS.md) preserve every original decision, composition, reading and inference task plus all 27 transfer tasks. [Diagnostics](DIAGNOSTICS.md) separately report probability quality, option-order sensitivity, missing evidence, native coverage and uncertainty. [Exact statistics](metrics/benchmark.json) include counts and confidence intervals.
32
 
33
+ ## Model and API scope
34
 
35
+ Decision models return typed Choice, Noul and Score answers in the SystemOne format. Encoder and decoder references are evaluated through their published native interfaces. Untuned Qwen models use the frozen letter-logit readout, with no generated-text parsing. Kev uses the matched BF16 backbone with FP32 head and shipped temperature; date-fact injection is disabled. This differs from the authors’ default FP32 reproduction. Laya English and multilingual are measured separately. Jev is a recorded hosted-service snapshot.
36
 
37
+ The tests assess decisions from the supplied state and fixed choices. They do not establish live fact retrieval, universally superior reasoning or cross-hardware speed rankings. [Immutable identities and evaluation provenance](metrics/evaluation-provenance.json).
MATERIALS.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "scope": "13model public comparison; presentation-only roster amendment",
3
+ "public_models": 13,
4
+ "task_rows": 54,
5
+ "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f",
6
+ "weights_runtime_temperature_unchanged": true,
7
+ "files": {
8
+ "ATTRIBUTIONS.md": "f13d3d134767e99988e721a35500a904ed0fdb72bb48aaf8519b06da3629473d",
9
+ "DIAGNOSTICS.md": "f0e3a5040f9843bf1480f0383d2f01f7e9a601347d6b1545d68bf861d1f64007",
10
+ "Dockerfile.runtime": "6c781fdb5ad2abfb8f61c906751ac449d01644c62704cbe28312fd7b7fb57521",
11
+ "EVALUATION.md": "0624b6aebb35e86e8456ed734027ae4590fbf24c91775e4e228fe89454780f14",
12
+ "LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
13
+ "METHODS.md": "8cc5a8edb9cad84736e19b712a9b9d141883ba7ead5aa3cdd45567180c1db7a4",
14
+ "QUESTION-SCALING.md": "2a04e4f06132108e022fdc694d9d66c4e30bbd5206149844649138f322e5ca18",
15
+ "QWEN-LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
16
+ "README.md": "e303076d1f0f3802a170c128d8a10221ddae6169cb0bef52f624811c9f79d5fd",
17
+ "RUNTIME.md": "91284e3281944e60330258fcbd2617905671da7330f02cac1b0217c987a30095",
18
+ "SENSITIVITY.md": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048",
19
+ "TASKS.md": "ffd8e2da2bed770855e69d54c44479bd66bf27a319684954ad8b55f742706341",
20
+ "USAGE.md": "49888ddf4ac8b87d60556a1f44083b4789c9c19d95de460514a26af1f264bc17",
21
+ "WEIGHTING.md": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048",
22
+ "assets/architecture.pdf": "482fac3bd270e5155e7d7fe61898c113279214222e4c2124b3e495b04b0fe1ea",
23
+ "assets/architecture.png": "1e8fbe7655e7c1aadb4b2323290cdb20e726c0a42b7914c0a1d5954a6a029931",
24
+ "assets/architecture.svg": "76a1f3165173ff07a4126947b9b6ae49da46f64ee2b1d740a3d9a35652fb6bae",
25
+ "assets/decision-matrix.pdf": "12f963382e810eb2adb8a9fbbc3b3b7e25e75e4c13d56a04db81419e0b565033",
26
+ "assets/decision-matrix.png": "29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4",
27
+ "assets/decision-matrix.svg": "95e90ed75b9d942e212b621a2968be3c5387d0bf637ca681c885062c5de93456",
28
+ "assets/decision-question-scaling.pdf": "d89a843d5b59b1bdd01124ec5e0733a41fa0cf8715ad9adf77cf8e525d749bc2",
29
+ "assets/decision-question-scaling.png": "fcab4ce4d6ae4876f9110d0231796926faa139a1fbfcd6c3e8ab4f44917be7e7",
30
+ "assets/decision-question-scaling.svg": "a88c94ace9d3acd6c6171943024d73cdee9e0301d5f4e7287bec488ddcf0b5db",
31
+ "assets/decision-ranking.pdf": "2c5d77b3a2420531e25a35e74389cd74968840117f90ab9f0f79e3e245101a97",
32
+ "assets/decision-ranking.png": "a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f",
33
+ "assets/decision-ranking.svg": "39be265fee1f5e6a0b051524fadc33786c0a923a28d771a774ea3c01ace460dc",
34
+ "assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
35
+ "assets/readout.svg": "701084bf7b0858a3adb247ae446df737b77d4b75004341dfd5f956b933f3d253",
36
+ "metrics/benchmark.json": "40b1da632cb586a5122fa34f85ea654e2890d6619bbe707b49afe829b4e8d3ce",
37
+ "metrics/evaluation-provenance.json": "cab27fc814933826d4a280ab280877c4683524a5ea67773e4e134a00fcf8befa",
38
+ "metrics/question-scaling.json": "859c474c4c1fb571c6041574698f6f47d4982fb56f1f2e68820d8a99b3a41337",
39
+ "model-card-example.json": "23915dd0232955a2345f53a4e716f43c3816e536a3e5418030db24687644de88",
40
+ "runtime-fla-requirements.lock": "35e1fedff9ca49092a7ada277bbd794abe7e8474e14a678a7e4bcec6406d0eb5"
41
+ }
42
+ }
README.md CHANGED
@@ -49,7 +49,6 @@ tags:
49
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
50
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
51
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
52
- | Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
53
 
54
  Accuracy (%), using the same five-panel decision benchmark. General decisions contribute 30%; composition contributes 25%; reading, inference and external transfer each contribute 15%. Bold marks a Decision model strictly above every external open reference in that column; Jev and other Decision models are excluded from this threshold. [Full tasks, uncertainty and comparator identities](EVALUATION.md).
55
 
 
49
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
50
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
51
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
 
52
 
53
  Accuracy (%), using the same five-panel decision benchmark. General decisions contribute 30%; composition contributes 25%; reading, inference and external transfer each contribute 15%. Bold marks a Decision model strictly above every external open reference in that column; Jev and other Decision models are excluded from this threshold. [Full tasks, uncertainty and comparator identities](EVALUATION.md).
54
 
SENSITIVITY.md CHANGED
@@ -15,4 +15,3 @@ Product weights were changed after earlier results were observed. These are iden
15
  | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
16
  | Laya · English | 51.03 | 50.85 | 51.76 |
17
  | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
18
- | Kai | 46.49 | 46.80 | 48.16 |
 
15
  | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
16
  | Laya · English | 51.03 | 50.85 | 51.76 |
17
  | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
 
TASKS.md CHANGED
@@ -5,93 +5,93 @@ Accuracy (%) on the same requested rows. Bold marks a Decision-family result str
5
  <details>
6
  <summary>Decisions · 10 tasks</summary>
7
 
8
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
9
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
10
- | News classification | 128 | 85.16 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 | 78.12 |
11
- | Boolean constraints | 64 | 100.00 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 | 53.12 |
12
- | Entity classification | 112 | 96.43 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 | 50.00 |
13
- | Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 | 81.25 |
14
- | Evidence placement | 96 | 30.21 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 | 1.04 |
15
- | Ordered rubric | 64 | 100.00 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 | 6.25 |
16
- | Relation composition | 96 | 56.25 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 | 51.04 |
17
- | Scoped evidence | 96 | 89.58 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 | 29.17 |
18
- | State tracking | 96 | 33.33 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 | 30.21 |
19
- | In / out of menu | 64 | 100.00 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 | 42.19 |
20
 
21
  </details>
22
 
23
  <details>
24
  <summary>Composition · 10 tasks</summary>
25
 
26
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
27
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
28
- | Record identity | 80 | 76.25 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 | 51.25 |
29
- | Capacity assignment | 80 | 68.75 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 | 50.00 |
30
- | Constraint assignment | 80 | 55.00 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 | 31.25 |
31
- | Intent routing · EN | 120 | 91.67 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 | 72.50 |
32
- | Intent routing · ZH | 120 | 88.33 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 | 67.50 |
33
- | Multiset reconciliation | 80 | 58.75 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 | 32.50 |
34
- | Ordered service loss | 80 | 42.50 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 | 21.25 |
35
- | Conflicting rule closure | 80 | 66.25 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 | 23.75 |
36
- | Temporal exclusion | 80 | 41.25 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 | 22.50 |
37
- | Transaction recovery | 80 | 75.00 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 | 26.25 |
38
 
39
  </details>
40
 
41
  <details>
42
  <summary>Reading · 3 tasks</summary>
43
 
44
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
45
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
46
- | Yes / no reading | 160 | 92.50 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 | 74.38 |
47
- | Reading · EN | 160 | 96.88 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 | 39.38 |
48
- | Reading · ZH | 160 | 96.25 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 | 35.62 |
49
 
50
  </details>
51
 
52
  <details>
53
  <summary>Inference · 4 tasks</summary>
54
 
55
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
56
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
57
- | Contextual reasoning | 120 | 86.67 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 | 32.50 |
58
- | Answerability | 120 | 90.83 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 | 55.83 |
59
- | Textual entailment | 120 | 82.50 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 | 34.17 |
60
- | Scientific inference | 120 | 99.17 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 | 95.83 |
61
 
62
  </details>
63
 
64
  <details>
65
  <summary>Transfer · 27 tasks</summary>
66
 
67
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
68
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
69
- | Buried emotion | 20 | 70.00 | 40.00 | 20.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 | 55.00 |
70
- | Buried paraphrase | 20 | 95.00 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 | 35.00 |
71
- | Buried entailment | 20 | 95.00 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 | 55.00 |
72
- | Buried offensive-language detection | 20 | 55.00 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 | 75.00 |
73
- | Combined policy conditions | 32 | 96.88 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 | 34.38 |
74
- | Policy exceptions | 32 | 100.00 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 | 62.50 |
75
- | Policy negation | 32 | 90.62 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 | 56.25 |
76
- | Authorization contrast | 40 | 100.00 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 | 50.00 |
77
- | Deadline contrast | 40 | 92.50 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 | 47.50 |
78
- | Emotion | 80 | 67.50 | 61.25 | 40.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 | 52.50 |
79
- | MMLU | 80 | 88.75 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 | 31.25 |
80
- | MMLU-Pro | 200 | 84.00 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 | 17.50 |
81
- | Paraphrase | 80 | 87.50 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 | 47.50 |
82
- | Question entailment | 80 | 91.25 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 | 71.25 |
83
- | Science questions | 80 | 100.00 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 | 92.50 |
84
- | Offensive-language detection | 80 | 76.25 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 | 78.75 |
85
- | Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 | 50.00 |
86
- | Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 | 50.00 |
87
- | Evidence control · deadline | 10 | 100.00 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 | 50.00 |
88
- | Evidence control · late fee | 10 | 100.00 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 | 30.00 |
89
- | Evidence control · quantity limit | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 | 50.00 |
90
- | Evidence control · return window | 10 | 50.00 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 | 50.00 |
91
- | Evidence control · shipping delay | 10 | 90.00 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 | 20.00 |
92
- | Evidence control · sla response | 10 | 100.00 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 | 20.00 |
93
- | Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 | 50.00 |
94
- | Evidence control · volume discount | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 | 30.00 |
95
- | Evidence control · warranty claim | 10 | 90.00 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 | 10.00 |
96
 
97
  </details>
 
5
  <details>
6
  <summary>Decisions · 10 tasks</summary>
7
 
8
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
9
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
10
+ | News classification | 128 | 85.16 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 |
11
+ | Boolean constraints | 64 | 100.00 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 |
12
+ | Entity classification | 112 | 96.43 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 |
13
+ | Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 |
14
+ | Evidence placement | 96 | 30.21 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 |
15
+ | Ordered rubric | 64 | 100.00 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 |
16
+ | Relation composition | 96 | 56.25 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 |
17
+ | Scoped evidence | 96 | 89.58 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 |
18
+ | State tracking | 96 | 33.33 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 |
19
+ | In / out of menu | 64 | 100.00 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 |
20
 
21
  </details>
22
 
23
  <details>
24
  <summary>Composition · 10 tasks</summary>
25
 
26
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
27
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
28
+ | Record identity | 80 | 76.25 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 |
29
+ | Capacity assignment | 80 | 68.75 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 |
30
+ | Constraint assignment | 80 | 55.00 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 |
31
+ | Intent routing · EN | 120 | 91.67 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 |
32
+ | Intent routing · ZH | 120 | 88.33 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 |
33
+ | Multiset reconciliation | 80 | 58.75 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 |
34
+ | Ordered service loss | 80 | 42.50 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 |
35
+ | Conflicting rule closure | 80 | 66.25 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 |
36
+ | Temporal exclusion | 80 | 41.25 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 |
37
+ | Transaction recovery | 80 | 75.00 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 |
38
 
39
  </details>
40
 
41
  <details>
42
  <summary>Reading · 3 tasks</summary>
43
 
44
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
45
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
46
+ | Yes / no reading | 160 | 92.50 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 |
47
+ | Reading · EN | 160 | 96.88 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 |
48
+ | Reading · ZH | 160 | 96.25 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 |
49
 
50
  </details>
51
 
52
  <details>
53
  <summary>Inference · 4 tasks</summary>
54
 
55
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
56
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
57
+ | Contextual reasoning | 120 | 86.67 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 |
58
+ | Answerability | 120 | 90.83 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 |
59
+ | Textual entailment | 120 | 82.50 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 |
60
+ | Scientific inference | 120 | 99.17 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 |
61
 
62
  </details>
63
 
64
  <details>
65
  <summary>Transfer · 27 tasks</summary>
66
 
67
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
68
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
69
+ | Buried emotion | 20 | 70.00 | 40.00 | 20.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 |
70
+ | Buried paraphrase | 20 | 95.00 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 |
71
+ | Buried entailment | 20 | 95.00 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 |
72
+ | Buried offensive-language detection | 20 | 55.00 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 |
73
+ | Combined policy conditions | 32 | 96.88 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 |
74
+ | Policy exceptions | 32 | 100.00 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 |
75
+ | Policy negation | 32 | 90.62 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 |
76
+ | Authorization contrast | 40 | 100.00 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 |
77
+ | Deadline contrast | 40 | 92.50 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 |
78
+ | Emotion | 80 | 67.50 | 61.25 | 40.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 |
79
+ | MMLU | 80 | 88.75 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 |
80
+ | MMLU-Pro | 200 | 84.00 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 |
81
+ | Paraphrase | 80 | 87.50 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 |
82
+ | Question entailment | 80 | 91.25 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 |
83
+ | Science questions | 80 | 100.00 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 |
84
+ | Offensive-language detection | 80 | 76.25 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 |
85
+ | Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 |
86
+ | Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 |
87
+ | Evidence control · deadline | 10 | 100.00 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 |
88
+ | Evidence control · late fee | 10 | 100.00 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 |
89
+ | Evidence control · quantity limit | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 |
90
+ | Evidence control · return window | 10 | 50.00 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 |
91
+ | Evidence control · shipping delay | 10 | 90.00 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 |
92
+ | Evidence control · sla response | 10 | 100.00 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 |
93
+ | Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 |
94
+ | Evidence control · volume discount | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 |
95
+ | Evidence control · warranty claim | 10 | 90.00 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 |
96
 
97
  </details>
WEIGHTING.md CHANGED
@@ -15,4 +15,3 @@ Product weights were changed after earlier results were observed. These are iden
15
  | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
16
  | Laya · English | 51.03 | 50.85 | 51.76 |
17
  | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
18
- | Kai | 46.49 | 46.80 | 48.16 |
 
15
  | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
16
  | Laya · English | 51.03 | 50.85 | 51.76 |
17
  | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
 
assets/decision-matrix.pdf CHANGED
Binary files a/assets/decision-matrix.pdf and b/assets/decision-matrix.pdf differ
 
assets/decision-matrix.png CHANGED

Git LFS Details

  • SHA256: b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc
  • Pointer size: 131 Bytes
  • Size of remote file: 345 kB

Git LFS Details

  • SHA256: 29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4
  • Pointer size: 131 Bytes
  • Size of remote file: 325 kB
assets/decision-matrix.svg CHANGED
assets/decision-ranking.pdf CHANGED
Binary files a/assets/decision-ranking.pdf and b/assets/decision-ranking.pdf differ
 
assets/decision-ranking.png CHANGED

Git LFS Details

  • SHA256: 1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0
  • Pointer size: 131 Bytes
  • Size of remote file: 226 kB

Git LFS Details

  • SHA256: a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f
  • Pointer size: 131 Bytes
  • Size of remote file: 215 kB
assets/decision-ranking.svg CHANGED
metrics/benchmark.json CHANGED
The diff for this file is too large to render. See raw diff
 
metrics/evaluation-provenance.json CHANGED
@@ -1,175 +1,4 @@
1
  {
2
- "protocol": {
3
- "version": "decision-priority-benchmark-v4",
4
- "utc": "2026-09-22T05:05:12.118255+00:00",
5
- "status": "User-requested product weighting, observed regression benchmark; release remains evidence-gated",
6
- "authorization": "Latest user explicitly requested transfer20%\u219215% and assigned released5% to the stronger original-decision/composition area. Original decisions chosen from already observed same-size gaps. This is outcome-informed product weighting, not neutral/blind prospective benchmark design.",
7
- "weights": {
8
- "old_core": "3/10",
9
- "v3_core": "1/4",
10
- "v4": "3/20",
11
- "v5": "3/20",
12
- "transfer_v9_test": "3/20"
13
- },
14
- "accuracy_mean": "Exact weighted categorical accuracy; originalwithin-panelweights and requested denominators unchanged.",
15
- "original_panel_registry_sha256": "1aefdf58d56e832e765194d0fa53a8647acea7d2ad2fc25ac95380bf85949fd2",
16
- "transfer_protocol_sha256": "f0c0f8982bfbb9edfc5e486f682906612fae579f28c58c2fa656ad9d10874d16",
17
- "transfer_payload_sha256": "97d82909bcbd8756904d6ed59b31e16d16086d62990ccdc981dc95b8c274e0c5",
18
- "transfer_labels_sha256": "0b64882d58e2ac1e0f5e2e102f916907554c94c07f2774b8c092683cfc81973c",
19
- "transfer_accuracy_questions": 1046,
20
- "original_core_questions": 2720,
21
- "diagnostics": "Preserve internal-v2 source-generalization/order/unknown/proper-score/efficiency axes separately; no millisecond or ECE averaging into accuracy.",
22
- "roster": {
23
- "Nox": {
24
- "role": "Decision family",
25
- "size": "4B",
26
- "repo": "llm-semantic-router/Decision-1.0-Nox"
27
- },
28
- "Sol": {
29
- "role": "Decision family",
30
- "size": "2B",
31
- "repo": "llm-semantic-router/Decision-1.0-Sol"
32
- },
33
- "Kai": {
34
- "role": "Decision family",
35
- "size": "0.572B encoder",
36
- "repo": "llm-semantic-router/Decision-1.0-Kai",
37
- "revision": "2079354070d02f4c2b1e60bce6b99f60dfad9622",
38
- "contract": "shipped1024complete-token admission; FP32; defaultB8; no limit override"
39
- },
40
- "Lux": {
41
- "role": "Decision family",
42
- "size": "9B",
43
- "repo": "llm-semantic-router/Decision-1.0-Lux",
44
- "pending_unreleased": true
45
- },
46
- "kev-0.8b": {
47
- "role": "open reference",
48
- "size": "0.8B",
49
- "repo": "jaredpalmer/kev-0.8b",
50
- "revision": "54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8"
51
- },
52
- "kev-4b": {
53
- "role": "open reference",
54
- "size": "4B",
55
- "repo": "jaredpalmer/kev-4b",
56
- "revision": "485ace8703592fcf405488b262449990824cfed1"
57
- },
58
- "kev-9b": {
59
- "role": "open reference",
60
- "size": "9B",
61
- "repo": "jaredpalmer/kev-9b",
62
- "revision": "2629c06a5aeb0feb3b9783bafed17ed8f39ecf5c"
63
- },
64
- "Decider": {
65
- "role": "open reference",
66
- "size": "2B",
67
- "repo": "Mapika/decider-2b"
68
- },
69
- "Laya-base": {
70
- "role": "open reference",
71
- "size": "0.421B encoder",
72
- "repo": "convaiinnovations/laya",
73
- "mode": "english"
74
- },
75
- "Laya-multilingual": {
76
- "role": "open reference",
77
- "size": "0.322B encoder",
78
- "repo": "convaiinnovations/laya",
79
- "mode": "multilingual"
80
- },
81
- "Jev": {
82
- "role": "closed frontier",
83
- "size": "undisclosed",
84
- "served_identity": "jev-1.13.0"
85
- },
86
- "Qwen3.5-2B": {
87
- "role": "untuned reference",
88
- "size": "2B",
89
- "repo": "Qwen/Qwen3.5-2B",
90
- "adapter": "original locked chat/LM-head letter readout; no training"
91
- },
92
- "Qwen3.5-4B": {
93
- "role": "untuned reference",
94
- "size": "4B",
95
- "repo": "Qwen/Qwen3.5-4B",
96
- "adapter": "original locked chat/LM-head letter readout; no training"
97
- },
98
- "Qwen3.5-9B": {
99
- "role": "untuned reference",
100
- "size": "9B",
101
- "repo": "Qwen/Qwen3.5-9B",
102
- "adapter": "original locked chat/LM-head letter readout; no training"
103
- }
104
- },
105
- "internal_extra_references": [
106
- "LLM2Jev2B",
107
- "LLM2Jev4B",
108
- "Nimble9B"
109
- ],
110
- "cohort_gates": {
111
- "Sol": {
112
- "required_same_size_open": [
113
- "Decider"
114
- ],
115
- "additional_measured_internal_peer": [
116
- "LLM2Jev2B"
117
- ],
118
- "required_same_size_untuned": [
119
- "Qwen3.5-2B"
120
- ]
121
- },
122
- "Nox": {
123
- "required_same_size_open": [
124
- "kev-4b"
125
- ],
126
- "additional_measured_internal_peer": [
127
- "LLM2Jev4B"
128
- ],
129
- "required_same_size_untuned": [
130
- "Qwen3.5-4B"
131
- ]
132
- },
133
- "Lux": {
134
- "must_exceed_current_family": [
135
- "Nox",
136
- "Sol"
137
- ],
138
- "required_same_size_open": [
139
- "kev-9b"
140
- ],
141
- "additional_measured_internal_peer": [
142
- "Nimble9B"
143
- ],
144
- "required_same_size_untuned": [
145
- "Qwen3.5-9B"
146
- ]
147
- }
148
- },
149
- "publication": {
150
- "first_expanded_cards": "Nox and Sol independently eligible only when their own required same-size comparisons are complete and strictly beaten; table includes complete available public roster. Unreleased Lux need not delay either card.",
151
- "later_updates": "Require newaggregate strictly higher than the current published same model plus maintained same-size lead; show regressions and full diagnostic metrics honestly.",
152
- "Lux_first_release": "Newweightedaggregate strictly exceeds then-currentNox/Sol and known same-sizeopen references; old4panelgate superseded before heldout/publication.",
153
- "missing_scores": "Never fill from different benchmark, revision, successful-only denominator or inferred model size.",
154
- "accuracy_uncertainty": "Report paired component bootstrap intervals; no new positive-CI gate invented.",
155
- "versions": "No trainingversion labels in visiblecard/table/figures; retain immutable revision/weight/adapter identities in machine-readable provenance.",
156
- "rank": "All public roster descending actual score. Decision family copper; distinct restrained colors for Kev, Decider, Laya, frontierJev and untunedQwen. White background, no logos or trainingversion labels.",
157
- "matrix": "Minimal readable typography, full task coverage, no logo; labels omit trainingversions.",
158
- "original_comparison": "Retain original four-panel and prior five-panel25/25/15/15/20 comparisons in methods/provenance. Explicit outcome-informed product-priority amendment; no universal or prospective performance claim."
159
- },
160
- "training_integrity": "Frozen ongoing training/SELECT/CAL unchanged. Testresults never select anothercheckpoint; future targetedtraining uses independent TRAIN/SELECT/CAL. Observed tests labeledregression.",
161
- "supersedes": {
162
- "protocol": {
163
- "path": "eval/decision-benchmark-v3/PROTOCOL.json",
164
- "sha256": "d0b7c667fd4de0dfa5268d2d83aebf40e2b0c43d615ad77f4c2264e35a137534"
165
- },
166
- "prior_full_statistics": {
167
- "path": "analysis/decision-benchmark-v3-results/full01/scored/STATISTICS.json",
168
- "sha256": "aaed8ca33a874300545b81566cb0ba37bf7cc8fcaac954beb9094aeef8ad82f7"
169
- }
170
- },
171
- "sensitivity": "Retain full25/25/15/15/20 results alongside30/25/15/15/15 in methods/provenance. Reweighting gains are never described as training improvements. No model selection, calibration, predictions or denominators changed."
172
- },
173
  "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
174
  "manifest_sha256": "3199a40b2719cad9b1e815629cf5ebad923a9a399188c1199a9d27e918bda9e3",
175
  "qualified_runtime": {
@@ -390,44 +219,6 @@
390
  }
391
  ]
392
  },
393
- "Kai": {
394
- "panels": {
395
- "old_core": {
396
- "path": "results/public-roster-baselines-v1/Kai/worker/old_core/normalized.jsonl",
397
- "sha256": "4758a99a75392b197c8902600a88edca65d6acea500e046e1ea8e4edd398a8a3"
398
- },
399
- "v3_core": {
400
- "path": "results/public-roster-baselines-v1/Kai/worker/v3_core/normalized.jsonl",
401
- "sha256": "7cdb285b88882a4c0d73c3da890056ef3b1d09e09bc01a893e73a12ff3bdbab8"
402
- },
403
- "v4": {
404
- "path": "results/public-roster-baselines-v1/Kai/worker/v4/normalized.jsonl",
405
- "sha256": "fa744ee2181fc4c379b748fdd4ee1080ffd877f43ebc9441d27a47013a2a9cec"
406
- },
407
- "v5": {
408
- "path": "results/public-roster-baselines-v1/Kai/worker/v5/normalized.jsonl",
409
- "sha256": "9d1e50d62733dbf06b824855938a9ee1311c030916cf173cc5b18fe066af48cd"
410
- }
411
- },
412
- "transfer_rows": {
413
- "path": "analysis/decision-benchmark-v3-baselines/Kai/ROWS.json",
414
- "sha256": "854950cb3888a49a801982341576f4a2ebdacbdd7e7a7e32d16f5777e0cbda46"
415
- },
416
- "evidence": [
417
- {
418
- "path": "results/public-roster-baselines-v1/Kai/worker/COMPLETE.json",
419
- "sha256": "39cc8bae7b59e02720123cd8efa9ba1939e02ef15a4bb3be89972b14eb7877cf"
420
- },
421
- {
422
- "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
423
- "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
424
- },
425
- {
426
- "path": "analysis/decision-benchmark-v3-baselines/Kai/REPORT.json",
427
- "sha256": "5ed1fd6a69d68fbd64bd45c4da9431a73fd6053574786ddd00323e0537fcefbc"
428
- }
429
- ]
430
- },
431
  "Laya-base": {
432
  "panels": {
433
  "v4": {
@@ -765,156 +556,6 @@
765
  "sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
766
  }
767
  ]
768
- },
769
- "llm2jev-2b": {
770
- "panels": {
771
- "old_core": {
772
- "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard0/llm2jev-2b/quality/normalized.jsonl",
773
- "sha256": "5f0aaf2be47d11c42ecb228aeb6f317bf2074c80d398a66969f06490d5847760"
774
- },
775
- "v3_core": {
776
- "path": "eval/heldout/v3/baselines/llm2jev-2b/core/normalized.jsonl",
777
- "sha256": "339cd97706404b3d1e5af9d07ae6d3efc79cbb40310cfcf1406c64477e24e460"
778
- },
779
- "v4": {
780
- "path": "results/open-reference-gap-v1/llm2jev-2b/v4/normalized.jsonl",
781
- "sha256": "98e6b3bc2113e18e00442a86b085e06981560f43cf6a8807c3d1818b3fd27caf"
782
- },
783
- "v5": {
784
- "path": "results/open-reference-gap-v1/llm2jev-2b/v5/normalized.jsonl",
785
- "sha256": "36b9d9e96bae101eb3db6600c412585f48551ff5908072963a5dd435632e4d65"
786
- }
787
- },
788
- "transfer_rows": {
789
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/ROWS.json",
790
- "sha256": "a37a6b19549896854fd0fb9bb32f594bb8cb60b840aa49ead0c004dcea00778e"
791
- },
792
- "evidence": [
793
- {
794
- "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/COMPLETE.json",
795
- "sha256": "ccd7e1f0d6a4193c629046961eb7171b1278a8dbae015ff216396ceac1f008ef"
796
- },
797
- {
798
- "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/METADATA.json",
799
- "sha256": "b0d0aa9840976e0330e676fdd67ce90230dbfec7502452956260188924583b2b"
800
- },
801
- {
802
- "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/predictions.jsonl",
803
- "sha256": "67a217eea9af720846f653d780973a64fff93111d8dcb9fa09b8b6f9594328a4"
804
- },
805
- {
806
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/REPORT.json",
807
- "sha256": "bb2bf635114fc805e6fef66dce70ac6ceca126008b3e8365871a6d12d9cd6e09"
808
- },
809
- {
810
- "path": "eval/open-reference-comparison-v1/REGISTRY.json",
811
- "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
812
- },
813
- {
814
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/COMPOSABLE-MANIFEST.json",
815
- "sha256": "071d3b47bd6a81e256e11be2ed2965c7416244bd512bd35be8baef5b8ebcb2be"
816
- }
817
- ]
818
- },
819
- "llm2jev-4b": {
820
- "panels": {
821
- "old_core": {
822
- "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/llm2jev-4b/quality/normalized.jsonl",
823
- "sha256": "8545cdc4ab9e9f870fdce1c1859632b8bac7505f6e961b4cc07e4f3b816328ce"
824
- },
825
- "v3_core": {
826
- "path": "eval/heldout/v3/baselines/llm2jev-4b/core/normalized.jsonl",
827
- "sha256": "b3bd38c90daf0bc3d77fea9a3ed0612049b07865bde5102356f82f7456002c97"
828
- },
829
- "v4": {
830
- "path": "results/open-reference-gap-v1/llm2jev-4b/v4/normalized.jsonl",
831
- "sha256": "409f8d2c57a1f8a4ef7dd18f3adb6cb8ee2ed544566206ccc909d54236ca95b7"
832
- },
833
- "v5": {
834
- "path": "results/open-reference-gap-v1/llm2jev-4b/v5/normalized.jsonl",
835
- "sha256": "d458e7e11cd2b1d66c7db6044408042630a48a474c9d3997141fc85dfb427f24"
836
- }
837
- },
838
- "transfer_rows": {
839
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/ROWS.json",
840
- "sha256": "abc8f7c04170544f5c4bef7a9af70ce0d7a25bfc4b4b8b1ba3e552413273bb7d"
841
- },
842
- "evidence": [
843
- {
844
- "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/COMPLETE.json",
845
- "sha256": "d7745d5ec1e08100629c55e2c7518ff408462e6adea4ebfb01b7b401b212c4b6"
846
- },
847
- {
848
- "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/METADATA.json",
849
- "sha256": "6a522387522b92a0c01cbd8f32cdaa154289318a98564297d2f5723bc4c3dee0"
850
- },
851
- {
852
- "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/predictions.jsonl",
853
- "sha256": "61bea4eb25014f6277cf81010872c1724e6d8f07a89803063649a21160b02edc"
854
- },
855
- {
856
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/REPORT.json",
857
- "sha256": "1097d4eeb4abc792cac569e4793c469a332dc24bbdf5568a242ce4cee4c03759"
858
- },
859
- {
860
- "path": "eval/open-reference-comparison-v1/REGISTRY.json",
861
- "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
862
- },
863
- {
864
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/COMPOSABLE-MANIFEST.json",
865
- "sha256": "82a7589e866610251b5b52f5adb20ab5c7ed9fa6fb2a184c9d6a3367576171e7"
866
- }
867
- ]
868
- },
869
- "nimble-9b": {
870
- "panels": {
871
- "old_core": {
872
- "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/nimble-9b/quality/normalized.jsonl",
873
- "sha256": "6fefc71c6fbea1106894c1081057c2c7a253fa4ebc2e6aecfe23abc347be08bb"
874
- },
875
- "v3_core": {
876
- "path": "eval/heldout/v3/baselines/nimble-9b/core/normalized.jsonl",
877
- "sha256": "ab5323762fc5d34961c0fbce78497627528b696c60e18b880d671b468db97bf9"
878
- },
879
- "v4": {
880
- "path": "results/open-reference-gap-v1/nimble-9b/v4/normalized.jsonl",
881
- "sha256": "6ae263cef9cafdbcd35cd1f1dcc5b1bd920d3e6023d4fc487970286c04ec8d05"
882
- },
883
- "v5": {
884
- "path": "eval/heldout/v5/open-baselines-v1/nimble-9b/core/normalized.jsonl",
885
- "sha256": "0fc6dae372265ef62ea87ce9dc07b63a77ccd33e3aa48e15754f5ab98c32580a"
886
- }
887
- },
888
- "transfer_rows": {
889
- "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/ROWS.json",
890
- "sha256": "f55f29f973b0974adc7fc3fd582e2e5c0414f440e4a094a7ec8e152685e62a28"
891
- },
892
- "evidence": [
893
- {
894
- "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/COMPLETE.json",
895
- "sha256": "efc5a1be536227db99ff9937b84e9a5bc688fc497c0e4a84d8261144cea8a51d"
896
- },
897
- {
898
- "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/METADATA.json",
899
- "sha256": "77eaa373ec6a74b483ac9b0cda739ecd8cce95a046266cd7a28488a768b82940"
900
- },
901
- {
902
- "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/predictions.jsonl",
903
- "sha256": "ad2212fa691c662c7917b33315439844339252778fdaefcc3080cdf46e00a434"
904
- },
905
- {
906
- "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/REPORT.json",
907
- "sha256": "b6065189b7522e5d0d2b7b088b673ba26212bee5fda2f0d7c6cd67fc9cc93136"
908
- },
909
- {
910
- "path": "eval/open-reference-comparison-v1/REGISTRY.json",
911
- "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
912
- },
913
- {
914
- "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/COMPOSABLE-MANIFEST.json",
915
- "sha256": "7423a2430984cf8e7eed2c5933ceeb0e2894d252e30e447d1294b6b6b08192bd"
916
- }
917
- ]
918
  }
919
  },
920
  "task_rows": 54,
@@ -931,8 +572,7 @@
931
  "kev-0.8b",
932
  "Qwen3.5-2B",
933
  "Laya-base",
934
- "Laya-multilingual",
935
- "Kai"
936
  ],
937
  "internal_extra_models_not_public_rank": [
938
  "llm2jev-2b",
@@ -1015,5 +655,7 @@
1015
  }
1016
  },
1017
  "source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
1018
- }
 
 
1019
  }
 
1
  {
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
  "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
3
  "manifest_sha256": "3199a40b2719cad9b1e815629cf5ebad923a9a399188c1199a9d27e918bda9e3",
4
  "qualified_runtime": {
 
219
  }
220
  ]
221
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
222
  "Laya-base": {
223
  "panels": {
224
  "v4": {
 
556
  "sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
557
  }
558
  ]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
559
  }
560
  },
561
  "task_rows": 54,
 
572
  "kev-0.8b",
573
  "Qwen3.5-2B",
574
  "Laya-base",
575
+ "Laya-multilingual"
 
576
  ],
577
  "internal_extra_models_not_public_rank": [
578
  "llm2jev-2b",
 
655
  }
656
  },
657
  "source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
658
+ },
659
+ "frozen_protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
660
+ "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
661
  }
release-manifest.json CHANGED
@@ -10,39 +10,41 @@
10
  "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
11
  "displayed_materials": {
12
  "ATTRIBUTIONS.md": "f13d3d134767e99988e721a35500a904ed0fdb72bb48aaf8519b06da3629473d",
13
- "DIAGNOSTICS.md": "5492ed49869e45e45650c2501d02a78f9a2fbc17267c54f90cf034a01ebda2b3",
14
  "Dockerfile.runtime": "6c781fdb5ad2abfb8f61c906751ac449d01644c62704cbe28312fd7b7fb57521",
15
- "EVALUATION.md": "7ea30f8dee24e3c05a52e28c07612fc0cfe158e122cf8a0553078558da28d87b",
16
  "LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
17
  "METHODS.md": "8cc5a8edb9cad84736e19b712a9b9d141883ba7ead5aa3cdd45567180c1db7a4",
18
  "QUESTION-SCALING.md": "2a04e4f06132108e022fdc694d9d66c4e30bbd5206149844649138f322e5ca18",
19
  "QWEN-LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
20
- "README.md": "128d35a759d269cfe88147c59b6335739155b3f86fb7182e5be6438bfc019ee4",
21
  "RUNTIME.md": "91284e3281944e60330258fcbd2617905671da7330f02cac1b0217c987a30095",
22
- "SENSITIVITY.md": "e03ec18140d5c18e606da23cc5b26ef41176ab38019912d85050677e4637f0a0",
23
- "TASKS.md": "e87c5271049f09ce23fcb7fdcbaa9f03e5c50039b6d7a4005a6b4c7fbc1d30ca",
24
  "USAGE.md": "49888ddf4ac8b87d60556a1f44083b4789c9c19d95de460514a26af1f264bc17",
25
- "WEIGHTING.md": "e03ec18140d5c18e606da23cc5b26ef41176ab38019912d85050677e4637f0a0",
26
  "assets/architecture.pdf": "482fac3bd270e5155e7d7fe61898c113279214222e4c2124b3e495b04b0fe1ea",
27
  "assets/architecture.png": "1e8fbe7655e7c1aadb4b2323290cdb20e726c0a42b7914c0a1d5954a6a029931",
28
  "assets/architecture.svg": "76a1f3165173ff07a4126947b9b6ae49da46f64ee2b1d740a3d9a35652fb6bae",
29
- "assets/decision-matrix.pdf": "d1651d91db893ffbdf3261ec476241eb03c18c7a769d19978c668d60a1ac844f",
30
- "assets/decision-matrix.png": "b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc",
31
- "assets/decision-matrix.svg": "dc92d1e3f3cf0518ee44f1bc512aae74732e1db82b919774a52c0da7c744b0b4",
32
  "assets/decision-question-scaling.pdf": "d89a843d5b59b1bdd01124ec5e0733a41fa0cf8715ad9adf77cf8e525d749bc2",
33
  "assets/decision-question-scaling.png": "fcab4ce4d6ae4876f9110d0231796926faa139a1fbfcd6c3e8ab4f44917be7e7",
34
  "assets/decision-question-scaling.svg": "a88c94ace9d3acd6c6171943024d73cdee9e0301d5f4e7287bec488ddcf0b5db",
35
- "assets/decision-ranking.pdf": "f3e037ada6c30871d9ebe0087fa3d122346f7ab1c927515878d524c492b7799a",
36
- "assets/decision-ranking.png": "1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0",
37
- "assets/decision-ranking.svg": "c8de982ba51fb9d2b729191018220964426a0f524a00a796553b8879db271d1a",
38
  "assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
39
  "assets/readout.svg": "701084bf7b0858a3adb247ae446df737b77d4b75004341dfd5f956b933f3d253",
40
- "metrics/benchmark.json": "50dac7884849cbdb8be58c1c82b811f51543a08a548253cb8295f4e6444c4718",
41
- "metrics/evaluation-provenance.json": "1f9c8b1e03d521160a8426798394418748daa72014dfc7272578d1e83fe237b4",
42
  "metrics/question-scaling.json": "859c474c4c1fb571c6041574698f6f47d4982fb56f1f2e68820d8a99b3a41337",
43
  "model-card-example.json": "23915dd0232955a2345f53a4e716f43c3816e536a3e5418030db24687644de88",
44
- "runtime-fla-requirements.lock": "35e1fedff9ca49092a7ada277bbd794abe7e8474e14a678a7e4bcec6406d0eb5"
 
45
  },
46
  "weight_identity": "all weights and runtime bound by bundle-manifest.json",
47
- "shared_diagnostics_materials_sha256": "47d511c7918112304b194e2797a7b0d9847312dacc705a0aca204b7bc53f4b0e"
 
48
  }
 
10
  "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
11
  "displayed_materials": {
12
  "ATTRIBUTIONS.md": "f13d3d134767e99988e721a35500a904ed0fdb72bb48aaf8519b06da3629473d",
13
+ "DIAGNOSTICS.md": "f0e3a5040f9843bf1480f0383d2f01f7e9a601347d6b1545d68bf861d1f64007",
14
  "Dockerfile.runtime": "6c781fdb5ad2abfb8f61c906751ac449d01644c62704cbe28312fd7b7fb57521",
15
+ "EVALUATION.md": "0624b6aebb35e86e8456ed734027ae4590fbf24c91775e4e228fe89454780f14",
16
  "LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
17
  "METHODS.md": "8cc5a8edb9cad84736e19b712a9b9d141883ba7ead5aa3cdd45567180c1db7a4",
18
  "QUESTION-SCALING.md": "2a04e4f06132108e022fdc694d9d66c4e30bbd5206149844649138f322e5ca18",
19
  "QWEN-LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
20
+ "README.md": "e303076d1f0f3802a170c128d8a10221ddae6169cb0bef52f624811c9f79d5fd",
21
  "RUNTIME.md": "91284e3281944e60330258fcbd2617905671da7330f02cac1b0217c987a30095",
22
+ "SENSITIVITY.md": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048",
23
+ "TASKS.md": "ffd8e2da2bed770855e69d54c44479bd66bf27a319684954ad8b55f742706341",
24
  "USAGE.md": "49888ddf4ac8b87d60556a1f44083b4789c9c19d95de460514a26af1f264bc17",
25
+ "WEIGHTING.md": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048",
26
  "assets/architecture.pdf": "482fac3bd270e5155e7d7fe61898c113279214222e4c2124b3e495b04b0fe1ea",
27
  "assets/architecture.png": "1e8fbe7655e7c1aadb4b2323290cdb20e726c0a42b7914c0a1d5954a6a029931",
28
  "assets/architecture.svg": "76a1f3165173ff07a4126947b9b6ae49da46f64ee2b1d740a3d9a35652fb6bae",
29
+ "assets/decision-matrix.pdf": "12f963382e810eb2adb8a9fbbc3b3b7e25e75e4c13d56a04db81419e0b565033",
30
+ "assets/decision-matrix.png": "29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4",
31
+ "assets/decision-matrix.svg": "95e90ed75b9d942e212b621a2968be3c5387d0bf637ca681c885062c5de93456",
32
  "assets/decision-question-scaling.pdf": "d89a843d5b59b1bdd01124ec5e0733a41fa0cf8715ad9adf77cf8e525d749bc2",
33
  "assets/decision-question-scaling.png": "fcab4ce4d6ae4876f9110d0231796926faa139a1fbfcd6c3e8ab4f44917be7e7",
34
  "assets/decision-question-scaling.svg": "a88c94ace9d3acd6c6171943024d73cdee9e0301d5f4e7287bec488ddcf0b5db",
35
+ "assets/decision-ranking.pdf": "2c5d77b3a2420531e25a35e74389cd74968840117f90ab9f0f79e3e245101a97",
36
+ "assets/decision-ranking.png": "a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f",
37
+ "assets/decision-ranking.svg": "39be265fee1f5e6a0b051524fadc33786c0a923a28d771a774ea3c01ace460dc",
38
  "assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
39
  "assets/readout.svg": "701084bf7b0858a3adb247ae446df737b77d4b75004341dfd5f956b933f3d253",
40
+ "metrics/benchmark.json": "40b1da632cb586a5122fa34f85ea654e2890d6619bbe707b49afe829b4e8d3ce",
41
+ "metrics/evaluation-provenance.json": "cab27fc814933826d4a280ab280877c4683524a5ea67773e4e134a00fcf8befa",
42
  "metrics/question-scaling.json": "859c474c4c1fb571c6041574698f6f47d4982fb56f1f2e68820d8a99b3a41337",
43
  "model-card-example.json": "23915dd0232955a2345f53a4e716f43c3816e536a3e5418030db24687644de88",
44
+ "runtime-fla-requirements.lock": "35e1fedff9ca49092a7ada277bbd794abe7e8474e14a678a7e4bcec6406d0eb5",
45
+ "MATERIALS.json": "cf00e48f3111ab392aa50fc2b8a0237d27c7aee6a2aaddd696c2af5d8d0cd493"
46
  },
47
  "weight_identity": "all weights and runtime bound by bundle-manifest.json",
48
+ "shared_diagnostics_materials_sha256": "cf00e48f3111ab392aa50fc2b8a0237d27c7aee6a2aaddd696c2af5d8d0cd493",
49
+ "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
50
  }