Xunzhuo commited on
Commit
46505c7
·
verified ·
1 Parent(s): 8aea799

Update public comparison roster; preserve model weights and benchmark scores

Browse files
DIAGNOSTICS.md CHANGED
@@ -19,7 +19,6 @@ These axes remain separate from headline accuracy. Probability metrics use the s
19
  | Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
20
  | Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
21
  | Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
22
- | Kai | 1046/1046 | 0.6066 | 1.1552 | 7.95 | 1.24 | 0.3343 |
23
 
24
  Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
25
 
@@ -40,7 +39,6 @@ Coverage at an error threshold keeps whole confidence-tie groups together. These
40
  | Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
41
  | Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
42
  | Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
43
- | Kai | 36/36 | 41.67 | 27.78 | 0.1170 |
44
 
45
  The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
46
 
@@ -61,7 +59,6 @@ The 36 paired permutations test the same semantics under changed option order. S
61
  | Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
62
  | Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
63
  | Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
64
- | Kai | 110/110 | 37.27 | 55.39 | 18.18 | 0.8262 | -0.07 |
65
 
66
  The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
67
 
@@ -82,9 +79,8 @@ The 110 unknowable examples have no scored true class and are excluded from accu
82
  | Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
83
  | Laya · English | 2720/2720 | 1264/1264 | 34 |
84
  | Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
85
- | Kai | 2720/2720 | 1264/1264 | 0 |
86
 
87
- Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation; Kai retains the shipped complete-request 1,024-token limit. Server-side truncation for Jev cannot be observed.
88
 
89
  ## Uncertainty
90
 
@@ -103,6 +99,5 @@ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and orde
103
  | Qwen3.5-2B | 57.24 | 55.74–58.76 |
104
  | Laya · English | 51.03 | 49.43–52.68 |
105
  | Laya · Multilingual | 47.19 | 45.58–48.82 |
106
- | Kai | 46.49 | 45.02–47.99 |
107
 
108
  Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
 
19
  | Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
20
  | Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
21
  | Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
 
22
 
23
  Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
24
 
 
39
  | Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
40
  | Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
41
  | Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
 
42
 
43
  The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
44
 
 
59
  | Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
60
  | Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
61
  | Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
 
62
 
63
  The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
64
 
 
79
  | Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
80
  | Laya · English | 2720/2720 | 1264/1264 | 34 |
81
  | Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
 
82
 
83
+ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation. Server-side truncation for Jev cannot be observed.
84
 
85
  ## Uncertainty
86
 
 
99
  | Qwen3.5-2B | 57.24 | 55.74–58.76 |
100
  | Laya · English | 51.03 | 49.43–52.68 |
101
  | Laya · Multilingual | 47.19 | 45.58–48.82 |
 
102
 
103
  Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
EVALUATION.md CHANGED
@@ -1,6 +1,6 @@
1
  # Evaluation
2
 
3
- The comparison covers **3,766 scored decisions across 54 tasks** and all 14 displayed models. All models use the same requested examples, candidate order and scoring rules. Each model keeps its native admission, truncation and probability semantics.
4
 
5
  | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
6
  |---|---:|---:|---:|---:|---:|---:|---:|
@@ -17,7 +17,6 @@ The comparison covers **3,766 scored decisions across 54 tasks** and all 14 disp
17
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
18
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
19
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
20
- | Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
21
 
22
  Accuracy (%). Rows are sorted by overall score. Bold marks Decision-family cells strictly above every external open or untuned reference in that metric; Jev and the other Decision models are excluded from that threshold.
23
 
@@ -33,6 +32,6 @@ These are **outcome-informed product-priority weights, chosen after observing be
33
 
34
  ## Model and API scope
35
 
36
- Decision models return typed Choice, Noul and Score answers in the SystemOne format. Encoder and decoder references are evaluated through their published native interfaces. Untuned Qwen models use the frozen letter-logit readout, with no generated-text parsing. Kev uses the matched BF16 backbone with FP32 head and shipped temperature; date-fact injection is disabled. This differs from the authors’ default FP32 reproduction. Laya English and multilingual are measured separately. Kai uses its shipped 1,024-token complete-request limit. Jev is a recorded hosted-service snapshot.
37
 
38
  The tests assess decisions from the supplied state and fixed choices. They do not establish live fact retrieval, universally superior reasoning or cross-hardware speed rankings. [Immutable identities and evaluation provenance](metrics/evaluation-provenance.json).
 
1
  # Evaluation
2
 
3
+ The comparison covers **3,766 scored decisions across 54 tasks** and all 13 displayed models. All models use the same requested examples, candidate order and scoring rules. Each model keeps its native admission, truncation and probability semantics.
4
 
5
  | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
6
  |---|---:|---:|---:|---:|---:|---:|---:|
 
17
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
18
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
19
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
 
20
 
21
  Accuracy (%). Rows are sorted by overall score. Bold marks Decision-family cells strictly above every external open or untuned reference in that metric; Jev and the other Decision models are excluded from that threshold.
22
 
 
32
 
33
  ## Model and API scope
34
 
35
+ Decision models return typed Choice, Noul and Score answers in the SystemOne format. Encoder and decoder references are evaluated through their published native interfaces. Untuned Qwen models use the frozen letter-logit readout, with no generated-text parsing. Kev uses the matched BF16 backbone with FP32 head and shipped temperature; date-fact injection is disabled. This differs from the authors’ default FP32 reproduction. Laya English and multilingual are measured separately. Jev is a recorded hosted-service snapshot.
36
 
37
  The tests assess decisions from the supplied state and fixed choices. They do not establish live fact retrieval, universally superior reasoning or cross-hardware speed rankings. [Immutable identities and evaluation provenance](metrics/evaluation-provenance.json).
MATERIALS.json CHANGED
@@ -1,36 +1,30 @@
1
  {
2
- "status": "reviewable-docs-only-payload",
3
- "family": "Nox",
4
- "gate_sha256": "687163d21bc76a3e096ade8860b98e55711692e7890c9f5fe263caba1352d3e9",
5
- "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
6
- "source_sha256": "92b0b306bfcb8f04c429687b6154695f6b058d168fed30eeae7c96c860bc991f",
7
- "quality_method": "outcome-informed product weights; unchanged model predictions",
8
- "all14_models": true,
9
- "all54_tasks": true,
10
- "latency_current_API_six_loads_30_samples": true,
11
- "model_or_runtime_changes": false,
12
- "public_secret_scan": "PASS",
13
  "files": {
14
- "DIAGNOSTICS.md": "5492ed49869e45e45650c2501d02a78f9a2fbc17267c54f90cf034a01ebda2b3",
15
- "EVALUATION.md": "3866dea6ae78b09bf93ad7b843d067960493217ae84b6ded87bfdac97b424608",
16
  "QUESTION-SCALING.md": "898f3554434284ff506dbc63c34859750f07b33dcf71f5b6d4ec7a798a56eb28",
17
- "README.md": "108a5727cda42f93dcc2a3d5fbe653cc42a5fccbe1574ad940553aafe2f1741d",
18
- "SENSITIVITY.md": "e03ec18140d5c18e606da23cc5b26ef41176ab38019912d85050677e4637f0a0",
19
- "TASKS.md": "e87c5271049f09ce23fcb7fdcbaa9f03e5c50039b6d7a4005a6b4c7fbc1d30ca",
20
  "assets/architecture.png": "8fcfa10b3941ee4aec1f5fba20a25e382a7d5de83074071c9ef68e627996de40",
21
  "assets/architecture.svg": "215d9af6f94cc24847fc1d23a0b2a287b36afdc1e10ea4f5c82895b2fe103002",
22
- "assets/decision-matrix.pdf": "d1651d91db893ffbdf3261ec476241eb03c18c7a769d19978c668d60a1ac844f",
23
- "assets/decision-matrix.png": "b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc",
24
- "assets/decision-matrix.svg": "dc92d1e3f3cf0518ee44f1bc512aae74732e1db82b919774a52c0da7c744b0b4",
25
  "assets/decision-question-scaling.pdf": "d756f60c142ecaf647e59ea5bb3a611a39d1572a6471cd49afb7ef602a23fb3b",
26
  "assets/decision-question-scaling.png": "722e664974b28324693f5f4af6609330c278385ac118ed6f757b8fc0afec64ac",
27
  "assets/decision-question-scaling.svg": "cc591eafa335b5d16edd306d6368679afd283a4b14d5f83fc77780a352b1d315",
28
- "assets/decision-ranking.pdf": "f3e037ada6c30871d9ebe0087fa3d122346f7ab1c927515878d524c492b7799a",
29
- "assets/decision-ranking.png": "1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0",
30
- "assets/decision-ranking.svg": "c8de982ba51fb9d2b729191018220964426a0f524a00a796553b8879db271d1a",
31
  "assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
32
- "metrics/benchmark.json": "50dac7884849cbdb8be58c1c82b811f51543a08a548253cb8295f4e6444c4718",
33
- "metrics/evaluation-provenance.json": "1f9c8b1e03d521160a8426798394418748daa72014dfc7272578d1e83fe237b4",
34
  "metrics/question-scaling.json": "9407f7ebf86581cdea21e704cdbebb9ced0a485eb980925f48233ebea0f830dc"
35
  }
36
  }
 
1
  {
2
+ "scope": "13model public comparison; presentation-only roster amendment",
3
+ "public_models": 13,
4
+ "task_rows": 54,
5
+ "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f",
6
+ "weights_runtime_temperature_unchanged": true,
 
 
 
 
 
 
7
  "files": {
8
+ "DIAGNOSTICS.md": "f0e3a5040f9843bf1480f0383d2f01f7e9a601347d6b1545d68bf861d1f64007",
9
+ "EVALUATION.md": "0624b6aebb35e86e8456ed734027ae4590fbf24c91775e4e228fe89454780f14",
10
  "QUESTION-SCALING.md": "898f3554434284ff506dbc63c34859750f07b33dcf71f5b6d4ec7a798a56eb28",
11
+ "README.md": "02678d253920bc54d7af864edbee273bb649a5268ede892a6602d6b023f66800",
12
+ "SENSITIVITY.md": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048",
13
+ "TASKS.md": "ffd8e2da2bed770855e69d54c44479bd66bf27a319684954ad8b55f742706341",
14
  "assets/architecture.png": "8fcfa10b3941ee4aec1f5fba20a25e382a7d5de83074071c9ef68e627996de40",
15
  "assets/architecture.svg": "215d9af6f94cc24847fc1d23a0b2a287b36afdc1e10ea4f5c82895b2fe103002",
16
+ "assets/decision-matrix.pdf": "12f963382e810eb2adb8a9fbbc3b3b7e25e75e4c13d56a04db81419e0b565033",
17
+ "assets/decision-matrix.png": "29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4",
18
+ "assets/decision-matrix.svg": "95e90ed75b9d942e212b621a2968be3c5387d0bf637ca681c885062c5de93456",
19
  "assets/decision-question-scaling.pdf": "d756f60c142ecaf647e59ea5bb3a611a39d1572a6471cd49afb7ef602a23fb3b",
20
  "assets/decision-question-scaling.png": "722e664974b28324693f5f4af6609330c278385ac118ed6f757b8fc0afec64ac",
21
  "assets/decision-question-scaling.svg": "cc591eafa335b5d16edd306d6368679afd283a4b14d5f83fc77780a352b1d315",
22
+ "assets/decision-ranking.pdf": "2c5d77b3a2420531e25a35e74389cd74968840117f90ab9f0f79e3e245101a97",
23
+ "assets/decision-ranking.png": "a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f",
24
+ "assets/decision-ranking.svg": "39be265fee1f5e6a0b051524fadc33786c0a923a28d771a774ea3c01ace460dc",
25
  "assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
26
+ "metrics/benchmark.json": "40b1da632cb586a5122fa34f85ea654e2890d6619bbe707b49afe829b4e8d3ce",
27
+ "metrics/evaluation-provenance.json": "cab27fc814933826d4a280ab280877c4683524a5ea67773e4e134a00fcf8befa",
28
  "metrics/question-scaling.json": "9407f7ebf86581cdea21e704cdbebb9ced0a485eb980925f48233ebea0f830dc"
29
  }
30
  }
README.md CHANGED
@@ -49,7 +49,6 @@ tags:
49
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
50
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
51
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
52
- | Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
53
 
54
  Accuracy (%). Overall weights: Decisions **30%**, Composition **25%**, Reading **15%**, Inference **15%**, Transfer **15%**. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.
55
 
 
49
  | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
50
  | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
51
  | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
 
52
 
53
  Accuracy (%). Overall weights: Decisions **30%**, Composition **25%**, Reading **15%**, Inference **15%**, Transfer **15%**. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.
54
 
SENSITIVITY.md CHANGED
@@ -15,4 +15,3 @@ Product weights were changed after earlier results were observed. These are iden
15
  | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
16
  | Laya · English | 51.03 | 50.85 | 51.76 |
17
  | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
18
- | Kai | 46.49 | 46.80 | 48.16 |
 
15
  | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
16
  | Laya · English | 51.03 | 50.85 | 51.76 |
17
  | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
 
TASKS.md CHANGED
@@ -5,93 +5,93 @@ Accuracy (%) on the same requested rows. Bold marks a Decision-family result str
5
  <details>
6
  <summary>Decisions · 10 tasks</summary>
7
 
8
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
9
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
10
- | News classification | 128 | 85.16 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 | 78.12 |
11
- | Boolean constraints | 64 | 100.00 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 | 53.12 |
12
- | Entity classification | 112 | 96.43 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 | 50.00 |
13
- | Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 | 81.25 |
14
- | Evidence placement | 96 | 30.21 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 | 1.04 |
15
- | Ordered rubric | 64 | 100.00 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 | 6.25 |
16
- | Relation composition | 96 | 56.25 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 | 51.04 |
17
- | Scoped evidence | 96 | 89.58 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 | 29.17 |
18
- | State tracking | 96 | 33.33 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 | 30.21 |
19
- | In / out of menu | 64 | 100.00 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 | 42.19 |
20
 
21
  </details>
22
 
23
  <details>
24
  <summary>Composition · 10 tasks</summary>
25
 
26
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
27
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
28
- | Record identity | 80 | 76.25 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 | 51.25 |
29
- | Capacity assignment | 80 | 68.75 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 | 50.00 |
30
- | Constraint assignment | 80 | 55.00 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 | 31.25 |
31
- | Intent routing · EN | 120 | 91.67 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 | 72.50 |
32
- | Intent routing · ZH | 120 | 88.33 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 | 67.50 |
33
- | Multiset reconciliation | 80 | 58.75 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 | 32.50 |
34
- | Ordered service loss | 80 | 42.50 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 | 21.25 |
35
- | Conflicting rule closure | 80 | 66.25 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 | 23.75 |
36
- | Temporal exclusion | 80 | 41.25 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 | 22.50 |
37
- | Transaction recovery | 80 | 75.00 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 | 26.25 |
38
 
39
  </details>
40
 
41
  <details>
42
  <summary>Reading · 3 tasks</summary>
43
 
44
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
45
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
46
- | Yes / no reading | 160 | 92.50 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 | 74.38 |
47
- | Reading · EN | 160 | 96.88 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 | 39.38 |
48
- | Reading · ZH | 160 | 96.25 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 | 35.62 |
49
 
50
  </details>
51
 
52
  <details>
53
  <summary>Inference · 4 tasks</summary>
54
 
55
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
56
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
57
- | Contextual reasoning | 120 | 86.67 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 | 32.50 |
58
- | Answerability | 120 | 90.83 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 | 55.83 |
59
- | Textual entailment | 120 | 82.50 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 | 34.17 |
60
- | Scientific inference | 120 | 99.17 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 | 95.83 |
61
 
62
  </details>
63
 
64
  <details>
65
  <summary>Transfer · 27 tasks</summary>
66
 
67
- | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
68
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
69
- | Buried emotion | 20 | 70.00 | 40.00 | 20.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 | 55.00 |
70
- | Buried paraphrase | 20 | 95.00 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 | 35.00 |
71
- | Buried entailment | 20 | 95.00 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 | 55.00 |
72
- | Buried offensive-language detection | 20 | 55.00 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 | 75.00 |
73
- | Combined policy conditions | 32 | 96.88 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 | 34.38 |
74
- | Policy exceptions | 32 | 100.00 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 | 62.50 |
75
- | Policy negation | 32 | 90.62 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 | 56.25 |
76
- | Authorization contrast | 40 | 100.00 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 | 50.00 |
77
- | Deadline contrast | 40 | 92.50 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 | 47.50 |
78
- | Emotion | 80 | 67.50 | 61.25 | 40.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 | 52.50 |
79
- | MMLU | 80 | 88.75 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 | 31.25 |
80
- | MMLU-Pro | 200 | 84.00 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 | 17.50 |
81
- | Paraphrase | 80 | 87.50 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 | 47.50 |
82
- | Question entailment | 80 | 91.25 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 | 71.25 |
83
- | Science questions | 80 | 100.00 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 | 92.50 |
84
- | Offensive-language detection | 80 | 76.25 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 | 78.75 |
85
- | Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 | 50.00 |
86
- | Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 | 50.00 |
87
- | Evidence control · deadline | 10 | 100.00 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 | 50.00 |
88
- | Evidence control · late fee | 10 | 100.00 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 | 30.00 |
89
- | Evidence control · quantity limit | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 | 50.00 |
90
- | Evidence control · return window | 10 | 50.00 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 | 50.00 |
91
- | Evidence control · shipping delay | 10 | 90.00 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 | 20.00 |
92
- | Evidence control · sla response | 10 | 100.00 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 | 20.00 |
93
- | Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 | 50.00 |
94
- | Evidence control · volume discount | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 | 30.00 |
95
- | Evidence control · warranty claim | 10 | 90.00 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 | 10.00 |
96
 
97
  </details>
 
5
  <details>
6
  <summary>Decisions · 10 tasks</summary>
7
 
8
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
9
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
10
+ | News classification | 128 | 85.16 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 |
11
+ | Boolean constraints | 64 | 100.00 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 |
12
+ | Entity classification | 112 | 96.43 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 |
13
+ | Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 |
14
+ | Evidence placement | 96 | 30.21 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 |
15
+ | Ordered rubric | 64 | 100.00 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 |
16
+ | Relation composition | 96 | 56.25 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 |
17
+ | Scoped evidence | 96 | 89.58 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 |
18
+ | State tracking | 96 | 33.33 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 |
19
+ | In / out of menu | 64 | 100.00 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 |
20
 
21
  </details>
22
 
23
  <details>
24
  <summary>Composition · 10 tasks</summary>
25
 
26
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
27
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
28
+ | Record identity | 80 | 76.25 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 |
29
+ | Capacity assignment | 80 | 68.75 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 |
30
+ | Constraint assignment | 80 | 55.00 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 |
31
+ | Intent routing · EN | 120 | 91.67 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 |
32
+ | Intent routing · ZH | 120 | 88.33 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 |
33
+ | Multiset reconciliation | 80 | 58.75 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 |
34
+ | Ordered service loss | 80 | 42.50 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 |
35
+ | Conflicting rule closure | 80 | 66.25 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 |
36
+ | Temporal exclusion | 80 | 41.25 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 |
37
+ | Transaction recovery | 80 | 75.00 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 |
38
 
39
  </details>
40
 
41
  <details>
42
  <summary>Reading · 3 tasks</summary>
43
 
44
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
45
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
46
+ | Yes / no reading | 160 | 92.50 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 |
47
+ | Reading · EN | 160 | 96.88 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 |
48
+ | Reading · ZH | 160 | 96.25 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 |
49
 
50
  </details>
51
 
52
  <details>
53
  <summary>Inference · 4 tasks</summary>
54
 
55
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
56
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
57
+ | Contextual reasoning | 120 | 86.67 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 |
58
+ | Answerability | 120 | 90.83 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 |
59
+ | Textual entailment | 120 | 82.50 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 |
60
+ | Scientific inference | 120 | 99.17 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 |
61
 
62
  </details>
63
 
64
  <details>
65
  <summary>Transfer · 27 tasks</summary>
66
 
67
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
68
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
69
+ | Buried emotion | 20 | 70.00 | 40.00 | 20.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 |
70
+ | Buried paraphrase | 20 | 95.00 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 |
71
+ | Buried entailment | 20 | 95.00 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 |
72
+ | Buried offensive-language detection | 20 | 55.00 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 |
73
+ | Combined policy conditions | 32 | 96.88 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 |
74
+ | Policy exceptions | 32 | 100.00 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 |
75
+ | Policy negation | 32 | 90.62 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 |
76
+ | Authorization contrast | 40 | 100.00 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 |
77
+ | Deadline contrast | 40 | 92.50 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 |
78
+ | Emotion | 80 | 67.50 | 61.25 | 40.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 |
79
+ | MMLU | 80 | 88.75 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 |
80
+ | MMLU-Pro | 200 | 84.00 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 |
81
+ | Paraphrase | 80 | 87.50 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 |
82
+ | Question entailment | 80 | 91.25 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 |
83
+ | Science questions | 80 | 100.00 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 |
84
+ | Offensive-language detection | 80 | 76.25 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 |
85
+ | Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 |
86
+ | Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 |
87
+ | Evidence control · deadline | 10 | 100.00 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 |
88
+ | Evidence control · late fee | 10 | 100.00 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 |
89
+ | Evidence control · quantity limit | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 |
90
+ | Evidence control · return window | 10 | 50.00 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 |
91
+ | Evidence control · shipping delay | 10 | 90.00 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 |
92
+ | Evidence control · sla response | 10 | 100.00 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 |
93
+ | Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 |
94
+ | Evidence control · volume discount | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 |
95
+ | Evidence control · warranty claim | 10 | 90.00 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 |
96
 
97
  </details>
assets/decision-matrix.pdf CHANGED
Binary files a/assets/decision-matrix.pdf and b/assets/decision-matrix.pdf differ
 
assets/decision-matrix.png CHANGED

Git LFS Details

  • SHA256: b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc
  • Pointer size: 131 Bytes
  • Size of remote file: 345 kB

Git LFS Details

  • SHA256: 29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4
  • Pointer size: 131 Bytes
  • Size of remote file: 325 kB
assets/decision-matrix.svg CHANGED
assets/decision-ranking.pdf CHANGED
Binary files a/assets/decision-ranking.pdf and b/assets/decision-ranking.pdf differ
 
assets/decision-ranking.png CHANGED

Git LFS Details

  • SHA256: 1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0
  • Pointer size: 131 Bytes
  • Size of remote file: 226 kB

Git LFS Details

  • SHA256: a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f
  • Pointer size: 131 Bytes
  • Size of remote file: 215 kB
assets/decision-ranking.svg CHANGED
metrics/benchmark.json CHANGED
The diff for this file is too large to render. See raw diff
 
metrics/evaluation-provenance.json CHANGED
@@ -1,175 +1,4 @@
1
  {
2
- "protocol": {
3
- "version": "decision-priority-benchmark-v4",
4
- "utc": "2026-09-22T05:05:12.118255+00:00",
5
- "status": "User-requested product weighting, observed regression benchmark; release remains evidence-gated",
6
- "authorization": "Latest user explicitly requested transfer20%\u219215% and assigned released5% to the stronger original-decision/composition area. Original decisions chosen from already observed same-size gaps. This is outcome-informed product weighting, not neutral/blind prospective benchmark design.",
7
- "weights": {
8
- "old_core": "3/10",
9
- "v3_core": "1/4",
10
- "v4": "3/20",
11
- "v5": "3/20",
12
- "transfer_v9_test": "3/20"
13
- },
14
- "accuracy_mean": "Exact weighted categorical accuracy; originalwithin-panelweights and requested denominators unchanged.",
15
- "original_panel_registry_sha256": "1aefdf58d56e832e765194d0fa53a8647acea7d2ad2fc25ac95380bf85949fd2",
16
- "transfer_protocol_sha256": "f0c0f8982bfbb9edfc5e486f682906612fae579f28c58c2fa656ad9d10874d16",
17
- "transfer_payload_sha256": "97d82909bcbd8756904d6ed59b31e16d16086d62990ccdc981dc95b8c274e0c5",
18
- "transfer_labels_sha256": "0b64882d58e2ac1e0f5e2e102f916907554c94c07f2774b8c092683cfc81973c",
19
- "transfer_accuracy_questions": 1046,
20
- "original_core_questions": 2720,
21
- "diagnostics": "Preserve internal-v2 source-generalization/order/unknown/proper-score/efficiency axes separately; no millisecond or ECE averaging into accuracy.",
22
- "roster": {
23
- "Nox": {
24
- "role": "Decision family",
25
- "size": "4B",
26
- "repo": "llm-semantic-router/Decision-1.0-Nox"
27
- },
28
- "Sol": {
29
- "role": "Decision family",
30
- "size": "2B",
31
- "repo": "llm-semantic-router/Decision-1.0-Sol"
32
- },
33
- "Kai": {
34
- "role": "Decision family",
35
- "size": "0.572B encoder",
36
- "repo": "llm-semantic-router/Decision-1.0-Kai",
37
- "revision": "2079354070d02f4c2b1e60bce6b99f60dfad9622",
38
- "contract": "shipped1024complete-token admission; FP32; defaultB8; no limit override"
39
- },
40
- "Lux": {
41
- "role": "Decision family",
42
- "size": "9B",
43
- "repo": "llm-semantic-router/Decision-1.0-Lux",
44
- "pending_unreleased": true
45
- },
46
- "kev-0.8b": {
47
- "role": "open reference",
48
- "size": "0.8B",
49
- "repo": "jaredpalmer/kev-0.8b",
50
- "revision": "54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8"
51
- },
52
- "kev-4b": {
53
- "role": "open reference",
54
- "size": "4B",
55
- "repo": "jaredpalmer/kev-4b",
56
- "revision": "485ace8703592fcf405488b262449990824cfed1"
57
- },
58
- "kev-9b": {
59
- "role": "open reference",
60
- "size": "9B",
61
- "repo": "jaredpalmer/kev-9b",
62
- "revision": "2629c06a5aeb0feb3b9783bafed17ed8f39ecf5c"
63
- },
64
- "Decider": {
65
- "role": "open reference",
66
- "size": "2B",
67
- "repo": "Mapika/decider-2b"
68
- },
69
- "Laya-base": {
70
- "role": "open reference",
71
- "size": "0.421B encoder",
72
- "repo": "convaiinnovations/laya",
73
- "mode": "english"
74
- },
75
- "Laya-multilingual": {
76
- "role": "open reference",
77
- "size": "0.322B encoder",
78
- "repo": "convaiinnovations/laya",
79
- "mode": "multilingual"
80
- },
81
- "Jev": {
82
- "role": "closed frontier",
83
- "size": "undisclosed",
84
- "served_identity": "jev-1.13.0"
85
- },
86
- "Qwen3.5-2B": {
87
- "role": "untuned reference",
88
- "size": "2B",
89
- "repo": "Qwen/Qwen3.5-2B",
90
- "adapter": "original locked chat/LM-head letter readout; no training"
91
- },
92
- "Qwen3.5-4B": {
93
- "role": "untuned reference",
94
- "size": "4B",
95
- "repo": "Qwen/Qwen3.5-4B",
96
- "adapter": "original locked chat/LM-head letter readout; no training"
97
- },
98
- "Qwen3.5-9B": {
99
- "role": "untuned reference",
100
- "size": "9B",
101
- "repo": "Qwen/Qwen3.5-9B",
102
- "adapter": "original locked chat/LM-head letter readout; no training"
103
- }
104
- },
105
- "internal_extra_references": [
106
- "LLM2Jev2B",
107
- "LLM2Jev4B",
108
- "Nimble9B"
109
- ],
110
- "cohort_gates": {
111
- "Sol": {
112
- "required_same_size_open": [
113
- "Decider"
114
- ],
115
- "additional_measured_internal_peer": [
116
- "LLM2Jev2B"
117
- ],
118
- "required_same_size_untuned": [
119
- "Qwen3.5-2B"
120
- ]
121
- },
122
- "Nox": {
123
- "required_same_size_open": [
124
- "kev-4b"
125
- ],
126
- "additional_measured_internal_peer": [
127
- "LLM2Jev4B"
128
- ],
129
- "required_same_size_untuned": [
130
- "Qwen3.5-4B"
131
- ]
132
- },
133
- "Lux": {
134
- "must_exceed_current_family": [
135
- "Nox",
136
- "Sol"
137
- ],
138
- "required_same_size_open": [
139
- "kev-9b"
140
- ],
141
- "additional_measured_internal_peer": [
142
- "Nimble9B"
143
- ],
144
- "required_same_size_untuned": [
145
- "Qwen3.5-9B"
146
- ]
147
- }
148
- },
149
- "publication": {
150
- "first_expanded_cards": "Nox and Sol independently eligible only when their own required same-size comparisons are complete and strictly beaten; table includes complete available public roster. Unreleased Lux need not delay either card.",
151
- "later_updates": "Require newaggregate strictly higher than the current published same model plus maintained same-size lead; show regressions and full diagnostic metrics honestly.",
152
- "Lux_first_release": "Newweightedaggregate strictly exceeds then-currentNox/Sol and known same-sizeopen references; old4panelgate superseded before heldout/publication.",
153
- "missing_scores": "Never fill from different benchmark, revision, successful-only denominator or inferred model size.",
154
- "accuracy_uncertainty": "Report paired component bootstrap intervals; no new positive-CI gate invented.",
155
- "versions": "No trainingversion labels in visiblecard/table/figures; retain immutable revision/weight/adapter identities in machine-readable provenance.",
156
- "rank": "All public roster descending actual score. Decision family copper; distinct restrained colors for Kev, Decider, Laya, frontierJev and untunedQwen. White background, no logos or trainingversion labels.",
157
- "matrix": "Minimal readable typography, full task coverage, no logo; labels omit trainingversions.",
158
- "original_comparison": "Retain original four-panel and prior five-panel25/25/15/15/20 comparisons in methods/provenance. Explicit outcome-informed product-priority amendment; no universal or prospective performance claim."
159
- },
160
- "training_integrity": "Frozen ongoing training/SELECT/CAL unchanged. Testresults never select anothercheckpoint; future targetedtraining uses independent TRAIN/SELECT/CAL. Observed tests labeledregression.",
161
- "supersedes": {
162
- "protocol": {
163
- "path": "eval/decision-benchmark-v3/PROTOCOL.json",
164
- "sha256": "d0b7c667fd4de0dfa5268d2d83aebf40e2b0c43d615ad77f4c2264e35a137534"
165
- },
166
- "prior_full_statistics": {
167
- "path": "analysis/decision-benchmark-v3-results/full01/scored/STATISTICS.json",
168
- "sha256": "aaed8ca33a874300545b81566cb0ba37bf7cc8fcaac954beb9094aeef8ad82f7"
169
- }
170
- },
171
- "sensitivity": "Retain full25/25/15/15/20 results alongside30/25/15/15/15 in methods/provenance. Reweighting gains are never described as training improvements. No model selection, calibration, predictions or denominators changed."
172
- },
173
  "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
174
  "manifest_sha256": "3199a40b2719cad9b1e815629cf5ebad923a9a399188c1199a9d27e918bda9e3",
175
  "qualified_runtime": {
@@ -390,44 +219,6 @@
390
  }
391
  ]
392
  },
393
- "Kai": {
394
- "panels": {
395
- "old_core": {
396
- "path": "results/public-roster-baselines-v1/Kai/worker/old_core/normalized.jsonl",
397
- "sha256": "4758a99a75392b197c8902600a88edca65d6acea500e046e1ea8e4edd398a8a3"
398
- },
399
- "v3_core": {
400
- "path": "results/public-roster-baselines-v1/Kai/worker/v3_core/normalized.jsonl",
401
- "sha256": "7cdb285b88882a4c0d73c3da890056ef3b1d09e09bc01a893e73a12ff3bdbab8"
402
- },
403
- "v4": {
404
- "path": "results/public-roster-baselines-v1/Kai/worker/v4/normalized.jsonl",
405
- "sha256": "fa744ee2181fc4c379b748fdd4ee1080ffd877f43ebc9441d27a47013a2a9cec"
406
- },
407
- "v5": {
408
- "path": "results/public-roster-baselines-v1/Kai/worker/v5/normalized.jsonl",
409
- "sha256": "9d1e50d62733dbf06b824855938a9ee1311c030916cf173cc5b18fe066af48cd"
410
- }
411
- },
412
- "transfer_rows": {
413
- "path": "analysis/decision-benchmark-v3-baselines/Kai/ROWS.json",
414
- "sha256": "854950cb3888a49a801982341576f4a2ebdacbdd7e7a7e32d16f5777e0cbda46"
415
- },
416
- "evidence": [
417
- {
418
- "path": "results/public-roster-baselines-v1/Kai/worker/COMPLETE.json",
419
- "sha256": "39cc8bae7b59e02720123cd8efa9ba1939e02ef15a4bb3be89972b14eb7877cf"
420
- },
421
- {
422
- "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
423
- "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
424
- },
425
- {
426
- "path": "analysis/decision-benchmark-v3-baselines/Kai/REPORT.json",
427
- "sha256": "5ed1fd6a69d68fbd64bd45c4da9431a73fd6053574786ddd00323e0537fcefbc"
428
- }
429
- ]
430
- },
431
  "Laya-base": {
432
  "panels": {
433
  "v4": {
@@ -765,156 +556,6 @@
765
  "sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
766
  }
767
  ]
768
- },
769
- "llm2jev-2b": {
770
- "panels": {
771
- "old_core": {
772
- "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard0/llm2jev-2b/quality/normalized.jsonl",
773
- "sha256": "5f0aaf2be47d11c42ecb228aeb6f317bf2074c80d398a66969f06490d5847760"
774
- },
775
- "v3_core": {
776
- "path": "eval/heldout/v3/baselines/llm2jev-2b/core/normalized.jsonl",
777
- "sha256": "339cd97706404b3d1e5af9d07ae6d3efc79cbb40310cfcf1406c64477e24e460"
778
- },
779
- "v4": {
780
- "path": "results/open-reference-gap-v1/llm2jev-2b/v4/normalized.jsonl",
781
- "sha256": "98e6b3bc2113e18e00442a86b085e06981560f43cf6a8807c3d1818b3fd27caf"
782
- },
783
- "v5": {
784
- "path": "results/open-reference-gap-v1/llm2jev-2b/v5/normalized.jsonl",
785
- "sha256": "36b9d9e96bae101eb3db6600c412585f48551ff5908072963a5dd435632e4d65"
786
- }
787
- },
788
- "transfer_rows": {
789
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/ROWS.json",
790
- "sha256": "a37a6b19549896854fd0fb9bb32f594bb8cb60b840aa49ead0c004dcea00778e"
791
- },
792
- "evidence": [
793
- {
794
- "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/COMPLETE.json",
795
- "sha256": "ccd7e1f0d6a4193c629046961eb7171b1278a8dbae015ff216396ceac1f008ef"
796
- },
797
- {
798
- "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/METADATA.json",
799
- "sha256": "b0d0aa9840976e0330e676fdd67ce90230dbfec7502452956260188924583b2b"
800
- },
801
- {
802
- "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/predictions.jsonl",
803
- "sha256": "67a217eea9af720846f653d780973a64fff93111d8dcb9fa09b8b6f9594328a4"
804
- },
805
- {
806
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/REPORT.json",
807
- "sha256": "bb2bf635114fc805e6fef66dce70ac6ceca126008b3e8365871a6d12d9cd6e09"
808
- },
809
- {
810
- "path": "eval/open-reference-comparison-v1/REGISTRY.json",
811
- "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
812
- },
813
- {
814
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/COMPOSABLE-MANIFEST.json",
815
- "sha256": "071d3b47bd6a81e256e11be2ed2965c7416244bd512bd35be8baef5b8ebcb2be"
816
- }
817
- ]
818
- },
819
- "llm2jev-4b": {
820
- "panels": {
821
- "old_core": {
822
- "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/llm2jev-4b/quality/normalized.jsonl",
823
- "sha256": "8545cdc4ab9e9f870fdce1c1859632b8bac7505f6e961b4cc07e4f3b816328ce"
824
- },
825
- "v3_core": {
826
- "path": "eval/heldout/v3/baselines/llm2jev-4b/core/normalized.jsonl",
827
- "sha256": "b3bd38c90daf0bc3d77fea9a3ed0612049b07865bde5102356f82f7456002c97"
828
- },
829
- "v4": {
830
- "path": "results/open-reference-gap-v1/llm2jev-4b/v4/normalized.jsonl",
831
- "sha256": "409f8d2c57a1f8a4ef7dd18f3adb6cb8ee2ed544566206ccc909d54236ca95b7"
832
- },
833
- "v5": {
834
- "path": "results/open-reference-gap-v1/llm2jev-4b/v5/normalized.jsonl",
835
- "sha256": "d458e7e11cd2b1d66c7db6044408042630a48a474c9d3997141fc85dfb427f24"
836
- }
837
- },
838
- "transfer_rows": {
839
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/ROWS.json",
840
- "sha256": "abc8f7c04170544f5c4bef7a9af70ce0d7a25bfc4b4b8b1ba3e552413273bb7d"
841
- },
842
- "evidence": [
843
- {
844
- "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/COMPLETE.json",
845
- "sha256": "d7745d5ec1e08100629c55e2c7518ff408462e6adea4ebfb01b7b401b212c4b6"
846
- },
847
- {
848
- "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/METADATA.json",
849
- "sha256": "6a522387522b92a0c01cbd8f32cdaa154289318a98564297d2f5723bc4c3dee0"
850
- },
851
- {
852
- "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/predictions.jsonl",
853
- "sha256": "61bea4eb25014f6277cf81010872c1724e6d8f07a89803063649a21160b02edc"
854
- },
855
- {
856
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/REPORT.json",
857
- "sha256": "1097d4eeb4abc792cac569e4793c469a332dc24bbdf5568a242ce4cee4c03759"
858
- },
859
- {
860
- "path": "eval/open-reference-comparison-v1/REGISTRY.json",
861
- "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
862
- },
863
- {
864
- "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/COMPOSABLE-MANIFEST.json",
865
- "sha256": "82a7589e866610251b5b52f5adb20ab5c7ed9fa6fb2a184c9d6a3367576171e7"
866
- }
867
- ]
868
- },
869
- "nimble-9b": {
870
- "panels": {
871
- "old_core": {
872
- "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/nimble-9b/quality/normalized.jsonl",
873
- "sha256": "6fefc71c6fbea1106894c1081057c2c7a253fa4ebc2e6aecfe23abc347be08bb"
874
- },
875
- "v3_core": {
876
- "path": "eval/heldout/v3/baselines/nimble-9b/core/normalized.jsonl",
877
- "sha256": "ab5323762fc5d34961c0fbce78497627528b696c60e18b880d671b468db97bf9"
878
- },
879
- "v4": {
880
- "path": "results/open-reference-gap-v1/nimble-9b/v4/normalized.jsonl",
881
- "sha256": "6ae263cef9cafdbcd35cd1f1dcc5b1bd920d3e6023d4fc487970286c04ec8d05"
882
- },
883
- "v5": {
884
- "path": "eval/heldout/v5/open-baselines-v1/nimble-9b/core/normalized.jsonl",
885
- "sha256": "0fc6dae372265ef62ea87ce9dc07b63a77ccd33e3aa48e15754f5ab98c32580a"
886
- }
887
- },
888
- "transfer_rows": {
889
- "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/ROWS.json",
890
- "sha256": "f55f29f973b0974adc7fc3fd582e2e5c0414f440e4a094a7ec8e152685e62a28"
891
- },
892
- "evidence": [
893
- {
894
- "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/COMPLETE.json",
895
- "sha256": "efc5a1be536227db99ff9937b84e9a5bc688fc497c0e4a84d8261144cea8a51d"
896
- },
897
- {
898
- "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/METADATA.json",
899
- "sha256": "77eaa373ec6a74b483ac9b0cda739ecd8cce95a046266cd7a28488a768b82940"
900
- },
901
- {
902
- "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/predictions.jsonl",
903
- "sha256": "ad2212fa691c662c7917b33315439844339252778fdaefcc3080cdf46e00a434"
904
- },
905
- {
906
- "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/REPORT.json",
907
- "sha256": "b6065189b7522e5d0d2b7b088b673ba26212bee5fda2f0d7c6cd67fc9cc93136"
908
- },
909
- {
910
- "path": "eval/open-reference-comparison-v1/REGISTRY.json",
911
- "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
912
- },
913
- {
914
- "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/COMPOSABLE-MANIFEST.json",
915
- "sha256": "7423a2430984cf8e7eed2c5933ceeb0e2894d252e30e447d1294b6b6b08192bd"
916
- }
917
- ]
918
  }
919
  },
920
  "task_rows": 54,
@@ -931,8 +572,7 @@
931
  "kev-0.8b",
932
  "Qwen3.5-2B",
933
  "Laya-base",
934
- "Laya-multilingual",
935
- "Kai"
936
  ],
937
  "internal_extra_models_not_public_rank": [
938
  "llm2jev-2b",
@@ -1015,5 +655,7 @@
1015
  }
1016
  },
1017
  "source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
1018
- }
 
 
1019
  }
 
1
  {
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
  "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
3
  "manifest_sha256": "3199a40b2719cad9b1e815629cf5ebad923a9a399188c1199a9d27e918bda9e3",
4
  "qualified_runtime": {
 
219
  }
220
  ]
221
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
222
  "Laya-base": {
223
  "panels": {
224
  "v4": {
 
556
  "sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
557
  }
558
  ]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
559
  }
560
  },
561
  "task_rows": 54,
 
572
  "kev-0.8b",
573
  "Qwen3.5-2B",
574
  "Laya-base",
575
+ "Laya-multilingual"
 
576
  ],
577
  "internal_extra_models_not_public_rank": [
578
  "llm2jev-2b",
 
655
  }
656
  },
657
  "source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
658
+ },
659
+ "frozen_protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
660
+ "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
661
  }
release-manifest.json CHANGED
@@ -3,7 +3,7 @@
3
  "status": "documentation-only-assembled",
4
  "bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
5
  "readiness_sha256": "fcbb84a9bf7294ac79e2421cdd2cbfdfc46a79fd74c41ba0965778472d1cafc1",
6
- "model_card_sha256": "108a5727cda42f93dcc2a3d5fbe653cc42a5fccbe1574ad940553aafe2f1741d",
7
  "repo_id": "llm-semantic-router/Decision-1.0-Nox",
8
  "assembly_script_sha256": "1697c4abaee22bb7b60259c042488178fc71847afe4810037683f874d7b8d513",
9
  "original_bundle_manifest_preserved": true,
@@ -16,8 +16,8 @@
16
  },
17
  {
18
  "file": "DIAGNOSTICS.md",
19
- "bytes": 5967,
20
- "sha256": "5492ed49869e45e45650c2501d02a78f9a2fbc17267c54f90cf034a01ebda2b3"
21
  },
22
  {
23
  "file": "Dockerfile.runtime",
@@ -26,8 +26,8 @@
26
  },
27
  {
28
  "file": "EVALUATION.md",
29
- "bytes": 3744,
30
- "sha256": "3866dea6ae78b09bf93ad7b843d067960493217ae84b6ded87bfdac97b424608"
31
  },
32
  {
33
  "file": "LICENSE",
@@ -36,8 +36,8 @@
36
  },
37
  {
38
  "file": "MATERIALS.json",
39
- "bytes": 2688,
40
- "sha256": "387ae91720084aefff2ce1964d30550485666587e7ea3a0a2bb5fae5af678d8d"
41
  },
42
  {
43
  "file": "NORMALIZATION_RUNTIME.md",
@@ -56,8 +56,8 @@
56
  },
57
  {
58
  "file": "README.md",
59
- "bytes": 4737,
60
- "sha256": "108a5727cda42f93dcc2a3d5fbe653cc42a5fccbe1574ad940553aafe2f1741d"
61
  },
62
  {
63
  "file": "RUNTIME-RELEASE.json",
@@ -76,8 +76,8 @@
76
  },
77
  {
78
  "file": "SENSITIVITY.md",
79
- "bytes": 783,
80
- "sha256": "e03ec18140d5c18e606da23cc5b26ef41176ab38019912d85050677e4637f0a0"
81
  },
82
  {
83
  "file": "SERVING_OPTIMIZATION.json",
@@ -91,8 +91,8 @@
91
  },
92
  {
93
  "file": "TASKS.md",
94
- "bytes": 9784,
95
- "sha256": "e87c5271049f09ce23fcb7fdcbaa9f03e5c50039b6d7a4005a6b4c7fbc1d30ca"
96
  },
97
  {
98
  "file": "USAGE.md",
@@ -236,18 +236,18 @@
236
  },
237
  {
238
  "file": "assets/decision-matrix.pdf",
239
- "bytes": 28634,
240
- "sha256": "d1651d91db893ffbdf3261ec476241eb03c18c7a769d19978c668d60a1ac844f"
241
  },
242
  {
243
  "file": "assets/decision-matrix.png",
244
- "bytes": 345190,
245
- "sha256": "b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc"
246
  },
247
  {
248
  "file": "assets/decision-matrix.svg",
249
- "bytes": 46970,
250
- "sha256": "dc92d1e3f3cf0518ee44f1bc512aae74732e1db82b919774a52c0da7c744b0b4"
251
  },
252
  {
253
  "file": "assets/decision-question-scaling-600px.png",
@@ -271,18 +271,18 @@
271
  },
272
  {
273
  "file": "assets/decision-ranking.pdf",
274
- "bytes": 24599,
275
- "sha256": "f3e037ada6c30871d9ebe0087fa3d122346f7ab1c927515878d524c492b7799a"
276
  },
277
  {
278
  "file": "assets/decision-ranking.png",
279
- "bytes": 226418,
280
- "sha256": "1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0"
281
  },
282
  {
283
  "file": "assets/decision-ranking.svg",
284
- "bytes": 14935,
285
- "sha256": "c8de982ba51fb9d2b729191018220964426a0f524a00a796553b8879db271d1a"
286
  },
287
  {
288
  "file": "assets/readout.png",
@@ -361,8 +361,8 @@
361
  },
362
  {
363
  "file": "metrics/benchmark.json",
364
- "bytes": 2451569,
365
- "sha256": "50dac7884849cbdb8be58c1c82b811f51543a08a548253cb8295f4e6444c4718"
366
  },
367
  {
368
  "file": "metrics/comparator-coverage.json",
@@ -371,8 +371,8 @@
371
  },
372
  {
373
  "file": "metrics/evaluation-provenance.json",
374
- "bytes": 44403,
375
- "sha256": "1f9c8b1e03d521160a8426798394418748daa72014dfc7272578d1e83fe237b4"
376
  },
377
  {
378
  "file": "metrics/expanded-quality.json",
@@ -557,9 +557,10 @@
557
  "MATERIALS.json": "copy",
558
  "metrics/semantic-consistency.json": "copy"
559
  },
560
- "scope": "Current14model five-panel product comparison,54task details, separate diagnostics,current API latency. Weights, runtime, tokenizer and calibration unchanged.",
561
  "release_tag": "v1.3.1",
562
- "change_kind": "documentation-only-current-benchmark",
563
- "previous_main_revision": "74979f7c7408325716dcb89b686d9cf94301923a",
564
- "previous_release_manifest_sha256": "7f54061edbe5e3801107a3792d07aaf94bf51e6e1b176b3e457c86b3477134ee"
 
565
  }
 
3
  "status": "documentation-only-assembled",
4
  "bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
5
  "readiness_sha256": "fcbb84a9bf7294ac79e2421cdd2cbfdfc46a79fd74c41ba0965778472d1cafc1",
6
+ "model_card_sha256": "02678d253920bc54d7af864edbee273bb649a5268ede892a6602d6b023f66800",
7
  "repo_id": "llm-semantic-router/Decision-1.0-Nox",
8
  "assembly_script_sha256": "1697c4abaee22bb7b60259c042488178fc71847afe4810037683f874d7b8d513",
9
  "original_bundle_manifest_preserved": true,
 
16
  },
17
  {
18
  "file": "DIAGNOSTICS.md",
19
+ "bytes": 5678,
20
+ "sha256": "f0e3a5040f9843bf1480f0383d2f01f7e9a601347d6b1545d68bf861d1f64007"
21
  },
22
  {
23
  "file": "Dockerfile.runtime",
 
26
  },
27
  {
28
  "file": "EVALUATION.md",
29
+ "bytes": 3622,
30
+ "sha256": "0624b6aebb35e86e8456ed734027ae4590fbf24c91775e4e228fe89454780f14"
31
  },
32
  {
33
  "file": "LICENSE",
 
36
  },
37
  {
38
  "file": "MATERIALS.json",
39
+ "bytes": 2379,
40
+ "sha256": "d52c25221dab8aa6ef1a3e8c91d6651a88c662246a00db4766f99af34e5a5271"
41
  },
42
  {
43
  "file": "NORMALIZATION_RUNTIME.md",
 
56
  },
57
  {
58
  "file": "README.md",
59
+ "bytes": 4672,
60
+ "sha256": "02678d253920bc54d7af864edbee273bb649a5268ede892a6602d6b023f66800"
61
  },
62
  {
63
  "file": "RUNTIME-RELEASE.json",
 
76
  },
77
  {
78
  "file": "SENSITIVITY.md",
79
+ "bytes": 751,
80
+ "sha256": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048"
81
  },
82
  {
83
  "file": "SERVING_OPTIMIZATION.json",
 
91
  },
92
  {
93
  "file": "TASKS.md",
94
+ "bytes": 9449,
95
+ "sha256": "ffd8e2da2bed770855e69d54c44479bd66bf27a319684954ad8b55f742706341"
96
  },
97
  {
98
  "file": "USAGE.md",
 
236
  },
237
  {
238
  "file": "assets/decision-matrix.pdf",
239
+ "bytes": 28175,
240
+ "sha256": "12f963382e810eb2adb8a9fbbc3b3b7e25e75e4c13d56a04db81419e0b565033"
241
  },
242
  {
243
  "file": "assets/decision-matrix.png",
244
+ "bytes": 324709,
245
+ "sha256": "29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4"
246
  },
247
  {
248
  "file": "assets/decision-matrix.svg",
249
+ "bytes": 43846,
250
+ "sha256": "95e90ed75b9d942e212b621a2968be3c5387d0bf637ca681c885062c5de93456"
251
  },
252
  {
253
  "file": "assets/decision-question-scaling-600px.png",
 
271
  },
272
  {
273
  "file": "assets/decision-ranking.pdf",
274
+ "bytes": 24211,
275
+ "sha256": "2c5d77b3a2420531e25a35e74389cd74968840117f90ab9f0f79e3e245101a97"
276
  },
277
  {
278
  "file": "assets/decision-ranking.png",
279
+ "bytes": 214771,
280
+ "sha256": "a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f"
281
  },
282
  {
283
  "file": "assets/decision-ranking.svg",
284
+ "bytes": 14180,
285
+ "sha256": "39be265fee1f5e6a0b051524fadc33786c0a923a28d771a774ea3c01ace460dc"
286
  },
287
  {
288
  "file": "assets/readout.png",
 
361
  },
362
  {
363
  "file": "metrics/benchmark.json",
364
+ "bytes": 2276797,
365
+ "sha256": "40b1da632cb586a5122fa34f85ea654e2890d6619bbe707b49afe829b4e8d3ce"
366
  },
367
  {
368
  "file": "metrics/comparator-coverage.json",
 
371
  },
372
  {
373
  "file": "metrics/evaluation-provenance.json",
374
+ "bytes": 28685,
375
+ "sha256": "cab27fc814933826d4a280ab280877c4683524a5ea67773e4e134a00fcf8befa"
376
  },
377
  {
378
  "file": "metrics/expanded-quality.json",
 
557
  "MATERIALS.json": "copy",
558
  "metrics/semantic-consistency.json": "copy"
559
  },
560
+ "scope": "13model public display only; original benchmark values and all model files unchanged.",
561
  "release_tag": "v1.3.1",
562
+ "change_kind": "documentation-only-public-roster",
563
+ "previous_main_revision": "8aea799ad37bf78ea8d3be42bde29c324524889e",
564
+ "previous_release_manifest_sha256": "752ff577b903fb71e1012d1b0a9ccafc83336e0a104126a7ff918389c3673198",
565
+ "presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
566
  }