Update public comparison roster; preserve model weights and benchmark scores
Browse files- DIAGNOSTICS.md +1 -6
- EVALUATION.md +10 -15
- MATERIALS.json +42 -0
- README.md +0 -1
- SENSITIVITY.md +0 -1
- TASKS.md +64 -64
- WEIGHTING.md +0 -1
- assets/decision-matrix.pdf +0 -0
- assets/decision-matrix.png +2 -2
- assets/decision-matrix.svg +612 -697
- assets/decision-ranking.pdf +0 -0
- assets/decision-ranking.png +2 -2
- assets/decision-ranking.svg +132 -149
- metrics/benchmark.json +0 -0
- metrics/evaluation-provenance.json +4 -362
- release-manifest.json +18 -16
DIAGNOSTICS.md
CHANGED
|
@@ -19,7 +19,6 @@ These axes remain separate from headline accuracy. Probability metrics use the s
|
|
| 19 |
| Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
|
| 20 |
| Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
|
| 21 |
| Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
|
| 22 |
-
| Kai | 1046/1046 | 0.6066 | 1.1552 | 7.95 | 1.24 | 0.3343 |
|
| 23 |
|
| 24 |
Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
|
| 25 |
|
|
@@ -40,7 +39,6 @@ Coverage at an error threshold keeps whole confidence-tie groups together. These
|
|
| 40 |
| Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
|
| 41 |
| Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
|
| 42 |
| Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
|
| 43 |
-
| Kai | 36/36 | 41.67 | 27.78 | 0.1170 |
|
| 44 |
|
| 45 |
The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
|
| 46 |
|
|
@@ -61,7 +59,6 @@ The 36 paired permutations test the same semantics under changed option order. S
|
|
| 61 |
| Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
|
| 62 |
| Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
|
| 63 |
| Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
|
| 64 |
-
| Kai | 110/110 | 37.27 | 55.39 | 18.18 | 0.8262 | -0.07 |
|
| 65 |
|
| 66 |
The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
|
| 67 |
|
|
@@ -82,9 +79,8 @@ The 110 unknowable examples have no scored true class and are excluded from accu
|
|
| 82 |
| Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
|
| 83 |
| Laya · English | 2720/2720 | 1264/1264 | 34 |
|
| 84 |
| Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
|
| 85 |
-
| Kai | 2720/2720 | 1264/1264 | 0 |
|
| 86 |
|
| 87 |
-
Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation
|
| 88 |
|
| 89 |
## Uncertainty
|
| 90 |
|
|
@@ -103,6 +99,5 @@ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and orde
|
|
| 103 |
| Qwen3.5-2B | 57.24 | 55.74–58.76 |
|
| 104 |
| Laya · English | 51.03 | 49.43–52.68 |
|
| 105 |
| Laya · Multilingual | 47.19 | 45.58–48.82 |
|
| 106 |
-
| Kai | 46.49 | 45.02–47.99 |
|
| 107 |
|
| 108 |
Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
|
|
|
|
| 19 |
| Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
|
| 20 |
| Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
|
| 21 |
| Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
|
|
|
|
| 22 |
|
| 23 |
Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
|
| 24 |
|
|
|
|
| 39 |
| Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
|
| 40 |
| Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
|
| 41 |
| Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
|
|
|
|
| 42 |
|
| 43 |
The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
|
| 44 |
|
|
|
|
| 59 |
| Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
|
| 60 |
| Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
|
| 61 |
| Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
|
|
|
|
| 62 |
|
| 63 |
The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
|
| 64 |
|
|
|
|
| 79 |
| Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
|
| 80 |
| Laya · English | 2720/2720 | 1264/1264 | 34 |
|
| 81 |
| Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
|
|
|
|
| 82 |
|
| 83 |
+
Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation. Server-side truncation for Jev cannot be observed.
|
| 84 |
|
| 85 |
## Uncertainty
|
| 86 |
|
|
|
|
| 99 |
| Qwen3.5-2B | 57.24 | 55.74–58.76 |
|
| 100 |
| Laya · English | 51.03 | 49.43–52.68 |
|
| 101 |
| Laya · Multilingual | 47.19 | 45.58–48.82 |
|
|
|
|
| 102 |
|
| 103 |
Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
|
EVALUATION.md
CHANGED
|
@@ -1,8 +1,6 @@
|
|
| 1 |
# Evaluation
|
| 2 |
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
**These product weights were chosen after earlier results were observed.** Reweighting changes the product score, not model predictions or trained capability. [Weight sensitivity](WEIGHTING.md) retains both the earlier five-panel weighting and the original four-panel comparison for these same current model predictions.
|
| 6 |
|
| 7 |
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|
| 8 |
|---|---:|---:|---:|---:|---:|---:|---:|
|
|
@@ -19,24 +17,21 @@ Lux achieves **76.72%** on the decision-focused benchmark. The headline combines
|
|
| 19 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
|
| 20 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
|
| 21 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
|
| 22 |
-
| Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
|
| 23 |
-
|
| 24 |
-
Models are ordered by unrounded overall accuracy. Bold identifies a Decision-family result strictly higher than every external open reference in that column, excluding other Decision-family models and the closed Jev frontier. Ties are not bold. Sizes identify the source-model tier; Lux's deployed text-plus-head parameter count is 7.941B. Kai is a 0.572B encoder. Laya English and multilingual are evaluated separately.
|
| 25 |
|
| 26 |
-
|
| 27 |
|
| 28 |
-
|
| 29 |
|
| 30 |
-
|
| 31 |
|
| 32 |
-
|
| 33 |
|
| 34 |
-
##
|
| 35 |
|
| 36 |
-
All
|
| 37 |
|
| 38 |
-
|
| 39 |
|
| 40 |
-
|
| 41 |
|
| 42 |
-
|
|
|
|
| 1 |
# Evaluation
|
| 2 |
|
| 3 |
+
The comparison covers **3,766 scored decisions across 54 tasks** and all 13 displayed models. All models use the same requested examples, candidate order and scoring rules. Each model keeps its native admission, truncation and probability semantics.
|
|
|
|
|
|
|
| 4 |
|
| 5 |
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|
| 6 |
|---|---:|---:|---:|---:|---:|---:|---:|
|
|
|
|
| 17 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
|
| 18 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
|
| 19 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
|
|
|
|
|
|
|
|
|
|
| 20 |
|
| 21 |
+
Accuracy (%). Rows are sorted by overall score. Bold marks Decision-family cells strictly above every external open or untuned reference in that metric; Jev and the other Decision models are excluded from that threshold.
|
| 22 |
|
| 23 |
+
## Scope and weighting
|
| 24 |
|
| 25 |
+
The overall score weights **Decisions 30%, Composition 25%, Reading 15%, Inference 15%, Transfer 15%**. The first four panels contain 880, 880, 480 and 480 questions and retain their original family/source weights. Transfer is micro-accuracy over 1,046 clean knowable questions from the frozen upstream transfer test; 110 missing-evidence questions and 108 variants remain separate diagnostics. No latency, calibration error or consistency score is averaged into accuracy.
|
| 26 |
|
| 27 |
+
These are **outcome-informed product-priority weights, chosen after observing benchmark results**. The data are observed regression tests, not a fresh blind test. Reweighting is not a training improvement. [Weight sensitivity](SENSITIVITY.md) retains the prior weighting and original four-panel comparison for the same model weights. Training, checkpoint selection and calibration do not use these test labels.
|
| 28 |
|
| 29 |
+
## Full results
|
| 30 |
|
| 31 |
+
[All 54 task rows](TASKS.md) preserve every original decision, composition, reading and inference task plus all 27 transfer tasks. [Diagnostics](DIAGNOSTICS.md) separately report probability quality, option-order sensitivity, missing evidence, native coverage and uncertainty. [Exact statistics](metrics/benchmark.json) include counts and confidence intervals.
|
| 32 |
|
| 33 |
+
## Model and API scope
|
| 34 |
|
| 35 |
+
Decision models return typed Choice, Noul and Score answers in the SystemOne format. Encoder and decoder references are evaluated through their published native interfaces. Untuned Qwen models use the frozen letter-logit readout, with no generated-text parsing. Kev uses the matched BF16 backbone with FP32 head and shipped temperature; date-fact injection is disabled. This differs from the authors’ default FP32 reproduction. Laya English and multilingual are measured separately. Jev is a recorded hosted-service snapshot.
|
| 36 |
|
| 37 |
+
The tests assess decisions from the supplied state and fixed choices. They do not establish live fact retrieval, universally superior reasoning or cross-hardware speed rankings. [Immutable identities and evaluation provenance](metrics/evaluation-provenance.json).
|
MATERIALS.json
ADDED
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"scope": "13model public comparison; presentation-only roster amendment",
|
| 3 |
+
"public_models": 13,
|
| 4 |
+
"task_rows": 54,
|
| 5 |
+
"presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f",
|
| 6 |
+
"weights_runtime_temperature_unchanged": true,
|
| 7 |
+
"files": {
|
| 8 |
+
"ATTRIBUTIONS.md": "f13d3d134767e99988e721a35500a904ed0fdb72bb48aaf8519b06da3629473d",
|
| 9 |
+
"DIAGNOSTICS.md": "f0e3a5040f9843bf1480f0383d2f01f7e9a601347d6b1545d68bf861d1f64007",
|
| 10 |
+
"Dockerfile.runtime": "6c781fdb5ad2abfb8f61c906751ac449d01644c62704cbe28312fd7b7fb57521",
|
| 11 |
+
"EVALUATION.md": "0624b6aebb35e86e8456ed734027ae4590fbf24c91775e4e228fe89454780f14",
|
| 12 |
+
"LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
|
| 13 |
+
"METHODS.md": "8cc5a8edb9cad84736e19b712a9b9d141883ba7ead5aa3cdd45567180c1db7a4",
|
| 14 |
+
"QUESTION-SCALING.md": "2a04e4f06132108e022fdc694d9d66c4e30bbd5206149844649138f322e5ca18",
|
| 15 |
+
"QWEN-LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
|
| 16 |
+
"README.md": "e303076d1f0f3802a170c128d8a10221ddae6169cb0bef52f624811c9f79d5fd",
|
| 17 |
+
"RUNTIME.md": "91284e3281944e60330258fcbd2617905671da7330f02cac1b0217c987a30095",
|
| 18 |
+
"SENSITIVITY.md": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048",
|
| 19 |
+
"TASKS.md": "ffd8e2da2bed770855e69d54c44479bd66bf27a319684954ad8b55f742706341",
|
| 20 |
+
"USAGE.md": "49888ddf4ac8b87d60556a1f44083b4789c9c19d95de460514a26af1f264bc17",
|
| 21 |
+
"WEIGHTING.md": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048",
|
| 22 |
+
"assets/architecture.pdf": "482fac3bd270e5155e7d7fe61898c113279214222e4c2124b3e495b04b0fe1ea",
|
| 23 |
+
"assets/architecture.png": "1e8fbe7655e7c1aadb4b2323290cdb20e726c0a42b7914c0a1d5954a6a029931",
|
| 24 |
+
"assets/architecture.svg": "76a1f3165173ff07a4126947b9b6ae49da46f64ee2b1d740a3d9a35652fb6bae",
|
| 25 |
+
"assets/decision-matrix.pdf": "12f963382e810eb2adb8a9fbbc3b3b7e25e75e4c13d56a04db81419e0b565033",
|
| 26 |
+
"assets/decision-matrix.png": "29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4",
|
| 27 |
+
"assets/decision-matrix.svg": "95e90ed75b9d942e212b621a2968be3c5387d0bf637ca681c885062c5de93456",
|
| 28 |
+
"assets/decision-question-scaling.pdf": "d89a843d5b59b1bdd01124ec5e0733a41fa0cf8715ad9adf77cf8e525d749bc2",
|
| 29 |
+
"assets/decision-question-scaling.png": "fcab4ce4d6ae4876f9110d0231796926faa139a1fbfcd6c3e8ab4f44917be7e7",
|
| 30 |
+
"assets/decision-question-scaling.svg": "a88c94ace9d3acd6c6171943024d73cdee9e0301d5f4e7287bec488ddcf0b5db",
|
| 31 |
+
"assets/decision-ranking.pdf": "2c5d77b3a2420531e25a35e74389cd74968840117f90ab9f0f79e3e245101a97",
|
| 32 |
+
"assets/decision-ranking.png": "a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f",
|
| 33 |
+
"assets/decision-ranking.svg": "39be265fee1f5e6a0b051524fadc33786c0a923a28d771a774ea3c01ace460dc",
|
| 34 |
+
"assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
|
| 35 |
+
"assets/readout.svg": "701084bf7b0858a3adb247ae446df737b77d4b75004341dfd5f956b933f3d253",
|
| 36 |
+
"metrics/benchmark.json": "40b1da632cb586a5122fa34f85ea654e2890d6619bbe707b49afe829b4e8d3ce",
|
| 37 |
+
"metrics/evaluation-provenance.json": "cab27fc814933826d4a280ab280877c4683524a5ea67773e4e134a00fcf8befa",
|
| 38 |
+
"metrics/question-scaling.json": "859c474c4c1fb571c6041574698f6f47d4982fb56f1f2e68820d8a99b3a41337",
|
| 39 |
+
"model-card-example.json": "23915dd0232955a2345f53a4e716f43c3816e536a3e5418030db24687644de88",
|
| 40 |
+
"runtime-fla-requirements.lock": "35e1fedff9ca49092a7ada277bbd794abe7e8474e14a678a7e4bcec6406d0eb5"
|
| 41 |
+
}
|
| 42 |
+
}
|
README.md
CHANGED
|
@@ -49,7 +49,6 @@ tags:
|
|
| 49 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
|
| 50 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
|
| 51 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
|
| 52 |
-
| Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
|
| 53 |
|
| 54 |
Accuracy (%), using the same five-panel decision benchmark. General decisions contribute 30%; composition contributes 25%; reading, inference and external transfer each contribute 15%. Bold marks a Decision model strictly above every external open reference in that column; Jev and other Decision models are excluded from this threshold. [Full tasks, uncertainty and comparator identities](EVALUATION.md).
|
| 55 |
|
|
|
|
| 49 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
|
| 50 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
|
| 51 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
|
|
|
|
| 52 |
|
| 53 |
Accuracy (%), using the same five-panel decision benchmark. General decisions contribute 30%; composition contributes 25%; reading, inference and external transfer each contribute 15%. Bold marks a Decision model strictly above every external open reference in that column; Jev and other Decision models are excluded from this threshold. [Full tasks, uncertainty and comparator identities](EVALUATION.md).
|
| 54 |
|
SENSITIVITY.md
CHANGED
|
@@ -15,4 +15,3 @@ Product weights were changed after earlier results were observed. These are iden
|
|
| 15 |
| Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
|
| 16 |
| Laya · English | 51.03 | 50.85 | 51.76 |
|
| 17 |
| Laya · Multilingual | 47.19 | 47.18 | 48.56 |
|
| 18 |
-
| Kai | 46.49 | 46.80 | 48.16 |
|
|
|
|
| 15 |
| Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
|
| 16 |
| Laya · English | 51.03 | 50.85 | 51.76 |
|
| 17 |
| Laya · Multilingual | 47.19 | 47.18 | 48.56 |
|
|
|
TASKS.md
CHANGED
|
@@ -5,93 +5,93 @@ Accuracy (%) on the same requested rows. Bold marks a Decision-family result str
|
|
| 5 |
<details>
|
| 6 |
<summary>Decisions · 10 tasks</summary>
|
| 7 |
|
| 8 |
-
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
|
| 9 |
-
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 10 |
-
| News classification | 128 | 85.16 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 |
|
| 11 |
-
| Boolean constraints | 64 | 100.00 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 |
|
| 12 |
-
| Entity classification | 112 | 96.43 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 |
|
| 13 |
-
| Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 |
|
| 14 |
-
| Evidence placement | 96 | 30.21 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 |
|
| 15 |
-
| Ordered rubric | 64 | 100.00 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 |
|
| 16 |
-
| Relation composition | 96 | 56.25 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 |
|
| 17 |
-
| Scoped evidence | 96 | 89.58 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 |
|
| 18 |
-
| State tracking | 96 | 33.33 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 |
|
| 19 |
-
| In / out of menu | 64 | 100.00 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 |
|
| 20 |
|
| 21 |
</details>
|
| 22 |
|
| 23 |
<details>
|
| 24 |
<summary>Composition · 10 tasks</summary>
|
| 25 |
|
| 26 |
-
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
|
| 27 |
-
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 28 |
-
| Record identity | 80 | 76.25 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 |
|
| 29 |
-
| Capacity assignment | 80 | 68.75 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 |
|
| 30 |
-
| Constraint assignment | 80 | 55.00 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 |
|
| 31 |
-
| Intent routing · EN | 120 | 91.67 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 |
|
| 32 |
-
| Intent routing · ZH | 120 | 88.33 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 |
|
| 33 |
-
| Multiset reconciliation | 80 | 58.75 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 |
|
| 34 |
-
| Ordered service loss | 80 | 42.50 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 |
|
| 35 |
-
| Conflicting rule closure | 80 | 66.25 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 |
|
| 36 |
-
| Temporal exclusion | 80 | 41.25 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 |
|
| 37 |
-
| Transaction recovery | 80 | 75.00 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 |
|
| 38 |
|
| 39 |
</details>
|
| 40 |
|
| 41 |
<details>
|
| 42 |
<summary>Reading · 3 tasks</summary>
|
| 43 |
|
| 44 |
-
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
|
| 45 |
-
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 46 |
-
| Yes / no reading | 160 | 92.50 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 |
|
| 47 |
-
| Reading · EN | 160 | 96.88 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 |
|
| 48 |
-
| Reading · ZH | 160 | 96.25 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 |
|
| 49 |
|
| 50 |
</details>
|
| 51 |
|
| 52 |
<details>
|
| 53 |
<summary>Inference · 4 tasks</summary>
|
| 54 |
|
| 55 |
-
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
|
| 56 |
-
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 57 |
-
| Contextual reasoning | 120 | 86.67 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 |
|
| 58 |
-
| Answerability | 120 | 90.83 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 |
|
| 59 |
-
| Textual entailment | 120 | 82.50 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 |
|
| 60 |
-
| Scientific inference | 120 | 99.17 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 |
|
| 61 |
|
| 62 |
</details>
|
| 63 |
|
| 64 |
<details>
|
| 65 |
<summary>Transfer · 27 tasks</summary>
|
| 66 |
|
| 67 |
-
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
|
| 68 |
-
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 69 |
-
| Buried emotion | 20 | 70.00 | 40.00 | 20.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 |
|
| 70 |
-
| Buried paraphrase | 20 | 95.00 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 |
|
| 71 |
-
| Buried entailment | 20 | 95.00 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 |
|
| 72 |
-
| Buried offensive-language detection | 20 | 55.00 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 |
|
| 73 |
-
| Combined policy conditions | 32 | 96.88 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 |
|
| 74 |
-
| Policy exceptions | 32 | 100.00 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 |
|
| 75 |
-
| Policy negation | 32 | 90.62 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 |
|
| 76 |
-
| Authorization contrast | 40 | 100.00 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 |
|
| 77 |
-
| Deadline contrast | 40 | 92.50 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 |
|
| 78 |
-
| Emotion | 80 | 67.50 | 61.25 | 40.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 |
|
| 79 |
-
| MMLU | 80 | 88.75 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 |
|
| 80 |
-
| MMLU-Pro | 200 | 84.00 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 |
|
| 81 |
-
| Paraphrase | 80 | 87.50 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 |
|
| 82 |
-
| Question entailment | 80 | 91.25 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 |
|
| 83 |
-
| Science questions | 80 | 100.00 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 |
|
| 84 |
-
| Offensive-language detection | 80 | 76.25 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 |
|
| 85 |
-
| Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 |
|
| 86 |
-
| Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 |
|
| 87 |
-
| Evidence control · deadline | 10 | 100.00 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 |
|
| 88 |
-
| Evidence control · late fee | 10 | 100.00 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 |
|
| 89 |
-
| Evidence control · quantity limit | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 |
|
| 90 |
-
| Evidence control · return window | 10 | 50.00 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 |
|
| 91 |
-
| Evidence control · shipping delay | 10 | 90.00 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 |
|
| 92 |
-
| Evidence control · sla response | 10 | 100.00 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 |
|
| 93 |
-
| Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 |
|
| 94 |
-
| Evidence control · volume discount | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 |
|
| 95 |
-
| Evidence control · warranty claim | 10 | 90.00 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 |
|
| 96 |
|
| 97 |
</details>
|
|
|
|
| 5 |
<details>
|
| 6 |
<summary>Decisions · 10 tasks</summary>
|
| 7 |
|
| 8 |
+
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
|
| 9 |
+
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
| 10 |
+
| News classification | 128 | 85.16 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 |
|
| 11 |
+
| Boolean constraints | 64 | 100.00 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 |
|
| 12 |
+
| Entity classification | 112 | 96.43 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 |
|
| 13 |
+
| Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 |
|
| 14 |
+
| Evidence placement | 96 | 30.21 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 |
|
| 15 |
+
| Ordered rubric | 64 | 100.00 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 |
|
| 16 |
+
| Relation composition | 96 | 56.25 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 |
|
| 17 |
+
| Scoped evidence | 96 | 89.58 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 |
|
| 18 |
+
| State tracking | 96 | 33.33 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 |
|
| 19 |
+
| In / out of menu | 64 | 100.00 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 |
|
| 20 |
|
| 21 |
</details>
|
| 22 |
|
| 23 |
<details>
|
| 24 |
<summary>Composition · 10 tasks</summary>
|
| 25 |
|
| 26 |
+
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
|
| 27 |
+
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
| 28 |
+
| Record identity | 80 | 76.25 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 |
|
| 29 |
+
| Capacity assignment | 80 | 68.75 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 |
|
| 30 |
+
| Constraint assignment | 80 | 55.00 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 |
|
| 31 |
+
| Intent routing · EN | 120 | 91.67 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 |
|
| 32 |
+
| Intent routing · ZH | 120 | 88.33 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 |
|
| 33 |
+
| Multiset reconciliation | 80 | 58.75 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 |
|
| 34 |
+
| Ordered service loss | 80 | 42.50 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 |
|
| 35 |
+
| Conflicting rule closure | 80 | 66.25 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 |
|
| 36 |
+
| Temporal exclusion | 80 | 41.25 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 |
|
| 37 |
+
| Transaction recovery | 80 | 75.00 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 |
|
| 38 |
|
| 39 |
</details>
|
| 40 |
|
| 41 |
<details>
|
| 42 |
<summary>Reading · 3 tasks</summary>
|
| 43 |
|
| 44 |
+
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
|
| 45 |
+
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
| 46 |
+
| Yes / no reading | 160 | 92.50 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 |
|
| 47 |
+
| Reading · EN | 160 | 96.88 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 |
|
| 48 |
+
| Reading · ZH | 160 | 96.25 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 |
|
| 49 |
|
| 50 |
</details>
|
| 51 |
|
| 52 |
<details>
|
| 53 |
<summary>Inference · 4 tasks</summary>
|
| 54 |
|
| 55 |
+
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
|
| 56 |
+
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
| 57 |
+
| Contextual reasoning | 120 | 86.67 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 |
|
| 58 |
+
| Answerability | 120 | 90.83 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 |
|
| 59 |
+
| Textual entailment | 120 | 82.50 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 |
|
| 60 |
+
| Scientific inference | 120 | 99.17 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 |
|
| 61 |
|
| 62 |
</details>
|
| 63 |
|
| 64 |
<details>
|
| 65 |
<summary>Transfer · 27 tasks</summary>
|
| 66 |
|
| 67 |
+
| Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual |
|
| 68 |
+
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
| 69 |
+
| Buried emotion | 20 | 70.00 | 40.00 | 20.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 |
|
| 70 |
+
| Buried paraphrase | 20 | 95.00 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 |
|
| 71 |
+
| Buried entailment | 20 | 95.00 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 |
|
| 72 |
+
| Buried offensive-language detection | 20 | 55.00 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 |
|
| 73 |
+
| Combined policy conditions | 32 | 96.88 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 |
|
| 74 |
+
| Policy exceptions | 32 | 100.00 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 |
|
| 75 |
+
| Policy negation | 32 | 90.62 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 |
|
| 76 |
+
| Authorization contrast | 40 | 100.00 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 |
|
| 77 |
+
| Deadline contrast | 40 | 92.50 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 |
|
| 78 |
+
| Emotion | 80 | 67.50 | 61.25 | 40.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 |
|
| 79 |
+
| MMLU | 80 | 88.75 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 |
|
| 80 |
+
| MMLU-Pro | 200 | 84.00 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 |
|
| 81 |
+
| Paraphrase | 80 | 87.50 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 |
|
| 82 |
+
| Question entailment | 80 | 91.25 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 |
|
| 83 |
+
| Science questions | 80 | 100.00 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 |
|
| 84 |
+
| Offensive-language detection | 80 | 76.25 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 |
|
| 85 |
+
| Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 |
|
| 86 |
+
| Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 |
|
| 87 |
+
| Evidence control · deadline | 10 | 100.00 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 |
|
| 88 |
+
| Evidence control · late fee | 10 | 100.00 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 |
|
| 89 |
+
| Evidence control · quantity limit | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 |
|
| 90 |
+
| Evidence control · return window | 10 | 50.00 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 |
|
| 91 |
+
| Evidence control · shipping delay | 10 | 90.00 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 |
|
| 92 |
+
| Evidence control · sla response | 10 | 100.00 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 |
|
| 93 |
+
| Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 |
|
| 94 |
+
| Evidence control · volume discount | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 |
|
| 95 |
+
| Evidence control · warranty claim | 10 | 90.00 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 |
|
| 96 |
|
| 97 |
</details>
|
WEIGHTING.md
CHANGED
|
@@ -15,4 +15,3 @@ Product weights were changed after earlier results were observed. These are iden
|
|
| 15 |
| Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
|
| 16 |
| Laya · English | 51.03 | 50.85 | 51.76 |
|
| 17 |
| Laya · Multilingual | 47.19 | 47.18 | 48.56 |
|
| 18 |
-
| Kai | 46.49 | 46.80 | 48.16 |
|
|
|
|
| 15 |
| Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
|
| 16 |
| Laya · English | 51.03 | 50.85 | 51.76 |
|
| 17 |
| Laya · Multilingual | 47.19 | 47.18 | 48.56 |
|
|
|
assets/decision-matrix.pdf
CHANGED
|
Binary files a/assets/decision-matrix.pdf and b/assets/decision-matrix.pdf differ
|
|
|
assets/decision-matrix.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
assets/decision-matrix.svg
CHANGED
|
|
|
|
assets/decision-ranking.pdf
CHANGED
|
Binary files a/assets/decision-ranking.pdf and b/assets/decision-ranking.pdf differ
|
|
|
assets/decision-ranking.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
assets/decision-ranking.svg
CHANGED
|
|
|
|
metrics/benchmark.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
metrics/evaluation-provenance.json
CHANGED
|
@@ -1,175 +1,4 @@
|
|
| 1 |
{
|
| 2 |
-
"protocol": {
|
| 3 |
-
"version": "decision-priority-benchmark-v4",
|
| 4 |
-
"utc": "2026-09-22T05:05:12.118255+00:00",
|
| 5 |
-
"status": "User-requested product weighting, observed regression benchmark; release remains evidence-gated",
|
| 6 |
-
"authorization": "Latest user explicitly requested transfer20%\u219215% and assigned released5% to the stronger original-decision/composition area. Original decisions chosen from already observed same-size gaps. This is outcome-informed product weighting, not neutral/blind prospective benchmark design.",
|
| 7 |
-
"weights": {
|
| 8 |
-
"old_core": "3/10",
|
| 9 |
-
"v3_core": "1/4",
|
| 10 |
-
"v4": "3/20",
|
| 11 |
-
"v5": "3/20",
|
| 12 |
-
"transfer_v9_test": "3/20"
|
| 13 |
-
},
|
| 14 |
-
"accuracy_mean": "Exact weighted categorical accuracy; originalwithin-panelweights and requested denominators unchanged.",
|
| 15 |
-
"original_panel_registry_sha256": "1aefdf58d56e832e765194d0fa53a8647acea7d2ad2fc25ac95380bf85949fd2",
|
| 16 |
-
"transfer_protocol_sha256": "f0c0f8982bfbb9edfc5e486f682906612fae579f28c58c2fa656ad9d10874d16",
|
| 17 |
-
"transfer_payload_sha256": "97d82909bcbd8756904d6ed59b31e16d16086d62990ccdc981dc95b8c274e0c5",
|
| 18 |
-
"transfer_labels_sha256": "0b64882d58e2ac1e0f5e2e102f916907554c94c07f2774b8c092683cfc81973c",
|
| 19 |
-
"transfer_accuracy_questions": 1046,
|
| 20 |
-
"original_core_questions": 2720,
|
| 21 |
-
"diagnostics": "Preserve internal-v2 source-generalization/order/unknown/proper-score/efficiency axes separately; no millisecond or ECE averaging into accuracy.",
|
| 22 |
-
"roster": {
|
| 23 |
-
"Nox": {
|
| 24 |
-
"role": "Decision family",
|
| 25 |
-
"size": "4B",
|
| 26 |
-
"repo": "llm-semantic-router/Decision-1.0-Nox"
|
| 27 |
-
},
|
| 28 |
-
"Sol": {
|
| 29 |
-
"role": "Decision family",
|
| 30 |
-
"size": "2B",
|
| 31 |
-
"repo": "llm-semantic-router/Decision-1.0-Sol"
|
| 32 |
-
},
|
| 33 |
-
"Kai": {
|
| 34 |
-
"role": "Decision family",
|
| 35 |
-
"size": "0.572B encoder",
|
| 36 |
-
"repo": "llm-semantic-router/Decision-1.0-Kai",
|
| 37 |
-
"revision": "2079354070d02f4c2b1e60bce6b99f60dfad9622",
|
| 38 |
-
"contract": "shipped1024complete-token admission; FP32; defaultB8; no limit override"
|
| 39 |
-
},
|
| 40 |
-
"Lux": {
|
| 41 |
-
"role": "Decision family",
|
| 42 |
-
"size": "9B",
|
| 43 |
-
"repo": "llm-semantic-router/Decision-1.0-Lux",
|
| 44 |
-
"pending_unreleased": true
|
| 45 |
-
},
|
| 46 |
-
"kev-0.8b": {
|
| 47 |
-
"role": "open reference",
|
| 48 |
-
"size": "0.8B",
|
| 49 |
-
"repo": "jaredpalmer/kev-0.8b",
|
| 50 |
-
"revision": "54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8"
|
| 51 |
-
},
|
| 52 |
-
"kev-4b": {
|
| 53 |
-
"role": "open reference",
|
| 54 |
-
"size": "4B",
|
| 55 |
-
"repo": "jaredpalmer/kev-4b",
|
| 56 |
-
"revision": "485ace8703592fcf405488b262449990824cfed1"
|
| 57 |
-
},
|
| 58 |
-
"kev-9b": {
|
| 59 |
-
"role": "open reference",
|
| 60 |
-
"size": "9B",
|
| 61 |
-
"repo": "jaredpalmer/kev-9b",
|
| 62 |
-
"revision": "2629c06a5aeb0feb3b9783bafed17ed8f39ecf5c"
|
| 63 |
-
},
|
| 64 |
-
"Decider": {
|
| 65 |
-
"role": "open reference",
|
| 66 |
-
"size": "2B",
|
| 67 |
-
"repo": "Mapika/decider-2b"
|
| 68 |
-
},
|
| 69 |
-
"Laya-base": {
|
| 70 |
-
"role": "open reference",
|
| 71 |
-
"size": "0.421B encoder",
|
| 72 |
-
"repo": "convaiinnovations/laya",
|
| 73 |
-
"mode": "english"
|
| 74 |
-
},
|
| 75 |
-
"Laya-multilingual": {
|
| 76 |
-
"role": "open reference",
|
| 77 |
-
"size": "0.322B encoder",
|
| 78 |
-
"repo": "convaiinnovations/laya",
|
| 79 |
-
"mode": "multilingual"
|
| 80 |
-
},
|
| 81 |
-
"Jev": {
|
| 82 |
-
"role": "closed frontier",
|
| 83 |
-
"size": "undisclosed",
|
| 84 |
-
"served_identity": "jev-1.13.0"
|
| 85 |
-
},
|
| 86 |
-
"Qwen3.5-2B": {
|
| 87 |
-
"role": "untuned reference",
|
| 88 |
-
"size": "2B",
|
| 89 |
-
"repo": "Qwen/Qwen3.5-2B",
|
| 90 |
-
"adapter": "original locked chat/LM-head letter readout; no training"
|
| 91 |
-
},
|
| 92 |
-
"Qwen3.5-4B": {
|
| 93 |
-
"role": "untuned reference",
|
| 94 |
-
"size": "4B",
|
| 95 |
-
"repo": "Qwen/Qwen3.5-4B",
|
| 96 |
-
"adapter": "original locked chat/LM-head letter readout; no training"
|
| 97 |
-
},
|
| 98 |
-
"Qwen3.5-9B": {
|
| 99 |
-
"role": "untuned reference",
|
| 100 |
-
"size": "9B",
|
| 101 |
-
"repo": "Qwen/Qwen3.5-9B",
|
| 102 |
-
"adapter": "original locked chat/LM-head letter readout; no training"
|
| 103 |
-
}
|
| 104 |
-
},
|
| 105 |
-
"internal_extra_references": [
|
| 106 |
-
"LLM2Jev2B",
|
| 107 |
-
"LLM2Jev4B",
|
| 108 |
-
"Nimble9B"
|
| 109 |
-
],
|
| 110 |
-
"cohort_gates": {
|
| 111 |
-
"Sol": {
|
| 112 |
-
"required_same_size_open": [
|
| 113 |
-
"Decider"
|
| 114 |
-
],
|
| 115 |
-
"additional_measured_internal_peer": [
|
| 116 |
-
"LLM2Jev2B"
|
| 117 |
-
],
|
| 118 |
-
"required_same_size_untuned": [
|
| 119 |
-
"Qwen3.5-2B"
|
| 120 |
-
]
|
| 121 |
-
},
|
| 122 |
-
"Nox": {
|
| 123 |
-
"required_same_size_open": [
|
| 124 |
-
"kev-4b"
|
| 125 |
-
],
|
| 126 |
-
"additional_measured_internal_peer": [
|
| 127 |
-
"LLM2Jev4B"
|
| 128 |
-
],
|
| 129 |
-
"required_same_size_untuned": [
|
| 130 |
-
"Qwen3.5-4B"
|
| 131 |
-
]
|
| 132 |
-
},
|
| 133 |
-
"Lux": {
|
| 134 |
-
"must_exceed_current_family": [
|
| 135 |
-
"Nox",
|
| 136 |
-
"Sol"
|
| 137 |
-
],
|
| 138 |
-
"required_same_size_open": [
|
| 139 |
-
"kev-9b"
|
| 140 |
-
],
|
| 141 |
-
"additional_measured_internal_peer": [
|
| 142 |
-
"Nimble9B"
|
| 143 |
-
],
|
| 144 |
-
"required_same_size_untuned": [
|
| 145 |
-
"Qwen3.5-9B"
|
| 146 |
-
]
|
| 147 |
-
}
|
| 148 |
-
},
|
| 149 |
-
"publication": {
|
| 150 |
-
"first_expanded_cards": "Nox and Sol independently eligible only when their own required same-size comparisons are complete and strictly beaten; table includes complete available public roster. Unreleased Lux need not delay either card.",
|
| 151 |
-
"later_updates": "Require newaggregate strictly higher than the current published same model plus maintained same-size lead; show regressions and full diagnostic metrics honestly.",
|
| 152 |
-
"Lux_first_release": "Newweightedaggregate strictly exceeds then-currentNox/Sol and known same-sizeopen references; old4panelgate superseded before heldout/publication.",
|
| 153 |
-
"missing_scores": "Never fill from different benchmark, revision, successful-only denominator or inferred model size.",
|
| 154 |
-
"accuracy_uncertainty": "Report paired component bootstrap intervals; no new positive-CI gate invented.",
|
| 155 |
-
"versions": "No trainingversion labels in visiblecard/table/figures; retain immutable revision/weight/adapter identities in machine-readable provenance.",
|
| 156 |
-
"rank": "All public roster descending actual score. Decision family copper; distinct restrained colors for Kev, Decider, Laya, frontierJev and untunedQwen. White background, no logos or trainingversion labels.",
|
| 157 |
-
"matrix": "Minimal readable typography, full task coverage, no logo; labels omit trainingversions.",
|
| 158 |
-
"original_comparison": "Retain original four-panel and prior five-panel25/25/15/15/20 comparisons in methods/provenance. Explicit outcome-informed product-priority amendment; no universal or prospective performance claim."
|
| 159 |
-
},
|
| 160 |
-
"training_integrity": "Frozen ongoing training/SELECT/CAL unchanged. Testresults never select anothercheckpoint; future targetedtraining uses independent TRAIN/SELECT/CAL. Observed tests labeledregression.",
|
| 161 |
-
"supersedes": {
|
| 162 |
-
"protocol": {
|
| 163 |
-
"path": "eval/decision-benchmark-v3/PROTOCOL.json",
|
| 164 |
-
"sha256": "d0b7c667fd4de0dfa5268d2d83aebf40e2b0c43d615ad77f4c2264e35a137534"
|
| 165 |
-
},
|
| 166 |
-
"prior_full_statistics": {
|
| 167 |
-
"path": "analysis/decision-benchmark-v3-results/full01/scored/STATISTICS.json",
|
| 168 |
-
"sha256": "aaed8ca33a874300545b81566cb0ba37bf7cc8fcaac954beb9094aeef8ad82f7"
|
| 169 |
-
}
|
| 170 |
-
},
|
| 171 |
-
"sensitivity": "Retain full25/25/15/15/20 results alongside30/25/15/15/15 in methods/provenance. Reweighting gains are never described as training improvements. No model selection, calibration, predictions or denominators changed."
|
| 172 |
-
},
|
| 173 |
"statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
|
| 174 |
"manifest_sha256": "3199a40b2719cad9b1e815629cf5ebad923a9a399188c1199a9d27e918bda9e3",
|
| 175 |
"qualified_runtime": {
|
|
@@ -390,44 +219,6 @@
|
|
| 390 |
}
|
| 391 |
]
|
| 392 |
},
|
| 393 |
-
"Kai": {
|
| 394 |
-
"panels": {
|
| 395 |
-
"old_core": {
|
| 396 |
-
"path": "results/public-roster-baselines-v1/Kai/worker/old_core/normalized.jsonl",
|
| 397 |
-
"sha256": "4758a99a75392b197c8902600a88edca65d6acea500e046e1ea8e4edd398a8a3"
|
| 398 |
-
},
|
| 399 |
-
"v3_core": {
|
| 400 |
-
"path": "results/public-roster-baselines-v1/Kai/worker/v3_core/normalized.jsonl",
|
| 401 |
-
"sha256": "7cdb285b88882a4c0d73c3da890056ef3b1d09e09bc01a893e73a12ff3bdbab8"
|
| 402 |
-
},
|
| 403 |
-
"v4": {
|
| 404 |
-
"path": "results/public-roster-baselines-v1/Kai/worker/v4/normalized.jsonl",
|
| 405 |
-
"sha256": "fa744ee2181fc4c379b748fdd4ee1080ffd877f43ebc9441d27a47013a2a9cec"
|
| 406 |
-
},
|
| 407 |
-
"v5": {
|
| 408 |
-
"path": "results/public-roster-baselines-v1/Kai/worker/v5/normalized.jsonl",
|
| 409 |
-
"sha256": "9d1e50d62733dbf06b824855938a9ee1311c030916cf173cc5b18fe066af48cd"
|
| 410 |
-
}
|
| 411 |
-
},
|
| 412 |
-
"transfer_rows": {
|
| 413 |
-
"path": "analysis/decision-benchmark-v3-baselines/Kai/ROWS.json",
|
| 414 |
-
"sha256": "854950cb3888a49a801982341576f4a2ebdacbdd7e7a7e32d16f5777e0cbda46"
|
| 415 |
-
},
|
| 416 |
-
"evidence": [
|
| 417 |
-
{
|
| 418 |
-
"path": "results/public-roster-baselines-v1/Kai/worker/COMPLETE.json",
|
| 419 |
-
"sha256": "39cc8bae7b59e02720123cd8efa9ba1939e02ef15a4bb3be89972b14eb7877cf"
|
| 420 |
-
},
|
| 421 |
-
{
|
| 422 |
-
"path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
|
| 423 |
-
"sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
|
| 424 |
-
},
|
| 425 |
-
{
|
| 426 |
-
"path": "analysis/decision-benchmark-v3-baselines/Kai/REPORT.json",
|
| 427 |
-
"sha256": "5ed1fd6a69d68fbd64bd45c4da9431a73fd6053574786ddd00323e0537fcefbc"
|
| 428 |
-
}
|
| 429 |
-
]
|
| 430 |
-
},
|
| 431 |
"Laya-base": {
|
| 432 |
"panels": {
|
| 433 |
"v4": {
|
|
@@ -765,156 +556,6 @@
|
|
| 765 |
"sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
|
| 766 |
}
|
| 767 |
]
|
| 768 |
-
},
|
| 769 |
-
"llm2jev-2b": {
|
| 770 |
-
"panels": {
|
| 771 |
-
"old_core": {
|
| 772 |
-
"path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard0/llm2jev-2b/quality/normalized.jsonl",
|
| 773 |
-
"sha256": "5f0aaf2be47d11c42ecb228aeb6f317bf2074c80d398a66969f06490d5847760"
|
| 774 |
-
},
|
| 775 |
-
"v3_core": {
|
| 776 |
-
"path": "eval/heldout/v3/baselines/llm2jev-2b/core/normalized.jsonl",
|
| 777 |
-
"sha256": "339cd97706404b3d1e5af9d07ae6d3efc79cbb40310cfcf1406c64477e24e460"
|
| 778 |
-
},
|
| 779 |
-
"v4": {
|
| 780 |
-
"path": "results/open-reference-gap-v1/llm2jev-2b/v4/normalized.jsonl",
|
| 781 |
-
"sha256": "98e6b3bc2113e18e00442a86b085e06981560f43cf6a8807c3d1818b3fd27caf"
|
| 782 |
-
},
|
| 783 |
-
"v5": {
|
| 784 |
-
"path": "results/open-reference-gap-v1/llm2jev-2b/v5/normalized.jsonl",
|
| 785 |
-
"sha256": "36b9d9e96bae101eb3db6600c412585f48551ff5908072963a5dd435632e4d65"
|
| 786 |
-
}
|
| 787 |
-
},
|
| 788 |
-
"transfer_rows": {
|
| 789 |
-
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/ROWS.json",
|
| 790 |
-
"sha256": "a37a6b19549896854fd0fb9bb32f594bb8cb60b840aa49ead0c004dcea00778e"
|
| 791 |
-
},
|
| 792 |
-
"evidence": [
|
| 793 |
-
{
|
| 794 |
-
"path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/COMPLETE.json",
|
| 795 |
-
"sha256": "ccd7e1f0d6a4193c629046961eb7171b1278a8dbae015ff216396ceac1f008ef"
|
| 796 |
-
},
|
| 797 |
-
{
|
| 798 |
-
"path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/METADATA.json",
|
| 799 |
-
"sha256": "b0d0aa9840976e0330e676fdd67ce90230dbfec7502452956260188924583b2b"
|
| 800 |
-
},
|
| 801 |
-
{
|
| 802 |
-
"path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/predictions.jsonl",
|
| 803 |
-
"sha256": "67a217eea9af720846f653d780973a64fff93111d8dcb9fa09b8b6f9594328a4"
|
| 804 |
-
},
|
| 805 |
-
{
|
| 806 |
-
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/REPORT.json",
|
| 807 |
-
"sha256": "bb2bf635114fc805e6fef66dce70ac6ceca126008b3e8365871a6d12d9cd6e09"
|
| 808 |
-
},
|
| 809 |
-
{
|
| 810 |
-
"path": "eval/open-reference-comparison-v1/REGISTRY.json",
|
| 811 |
-
"sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
|
| 812 |
-
},
|
| 813 |
-
{
|
| 814 |
-
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/COMPOSABLE-MANIFEST.json",
|
| 815 |
-
"sha256": "071d3b47bd6a81e256e11be2ed2965c7416244bd512bd35be8baef5b8ebcb2be"
|
| 816 |
-
}
|
| 817 |
-
]
|
| 818 |
-
},
|
| 819 |
-
"llm2jev-4b": {
|
| 820 |
-
"panels": {
|
| 821 |
-
"old_core": {
|
| 822 |
-
"path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/llm2jev-4b/quality/normalized.jsonl",
|
| 823 |
-
"sha256": "8545cdc4ab9e9f870fdce1c1859632b8bac7505f6e961b4cc07e4f3b816328ce"
|
| 824 |
-
},
|
| 825 |
-
"v3_core": {
|
| 826 |
-
"path": "eval/heldout/v3/baselines/llm2jev-4b/core/normalized.jsonl",
|
| 827 |
-
"sha256": "b3bd38c90daf0bc3d77fea9a3ed0612049b07865bde5102356f82f7456002c97"
|
| 828 |
-
},
|
| 829 |
-
"v4": {
|
| 830 |
-
"path": "results/open-reference-gap-v1/llm2jev-4b/v4/normalized.jsonl",
|
| 831 |
-
"sha256": "409f8d2c57a1f8a4ef7dd18f3adb6cb8ee2ed544566206ccc909d54236ca95b7"
|
| 832 |
-
},
|
| 833 |
-
"v5": {
|
| 834 |
-
"path": "results/open-reference-gap-v1/llm2jev-4b/v5/normalized.jsonl",
|
| 835 |
-
"sha256": "d458e7e11cd2b1d66c7db6044408042630a48a474c9d3997141fc85dfb427f24"
|
| 836 |
-
}
|
| 837 |
-
},
|
| 838 |
-
"transfer_rows": {
|
| 839 |
-
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/ROWS.json",
|
| 840 |
-
"sha256": "abc8f7c04170544f5c4bef7a9af70ce0d7a25bfc4b4b8b1ba3e552413273bb7d"
|
| 841 |
-
},
|
| 842 |
-
"evidence": [
|
| 843 |
-
{
|
| 844 |
-
"path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/COMPLETE.json",
|
| 845 |
-
"sha256": "d7745d5ec1e08100629c55e2c7518ff408462e6adea4ebfb01b7b401b212c4b6"
|
| 846 |
-
},
|
| 847 |
-
{
|
| 848 |
-
"path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/METADATA.json",
|
| 849 |
-
"sha256": "6a522387522b92a0c01cbd8f32cdaa154289318a98564297d2f5723bc4c3dee0"
|
| 850 |
-
},
|
| 851 |
-
{
|
| 852 |
-
"path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/predictions.jsonl",
|
| 853 |
-
"sha256": "61bea4eb25014f6277cf81010872c1724e6d8f07a89803063649a21160b02edc"
|
| 854 |
-
},
|
| 855 |
-
{
|
| 856 |
-
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/REPORT.json",
|
| 857 |
-
"sha256": "1097d4eeb4abc792cac569e4793c469a332dc24bbdf5568a242ce4cee4c03759"
|
| 858 |
-
},
|
| 859 |
-
{
|
| 860 |
-
"path": "eval/open-reference-comparison-v1/REGISTRY.json",
|
| 861 |
-
"sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
|
| 862 |
-
},
|
| 863 |
-
{
|
| 864 |
-
"path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/COMPOSABLE-MANIFEST.json",
|
| 865 |
-
"sha256": "82a7589e866610251b5b52f5adb20ab5c7ed9fa6fb2a184c9d6a3367576171e7"
|
| 866 |
-
}
|
| 867 |
-
]
|
| 868 |
-
},
|
| 869 |
-
"nimble-9b": {
|
| 870 |
-
"panels": {
|
| 871 |
-
"old_core": {
|
| 872 |
-
"path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/nimble-9b/quality/normalized.jsonl",
|
| 873 |
-
"sha256": "6fefc71c6fbea1106894c1081057c2c7a253fa4ebc2e6aecfe23abc347be08bb"
|
| 874 |
-
},
|
| 875 |
-
"v3_core": {
|
| 876 |
-
"path": "eval/heldout/v3/baselines/nimble-9b/core/normalized.jsonl",
|
| 877 |
-
"sha256": "ab5323762fc5d34961c0fbce78497627528b696c60e18b880d671b468db97bf9"
|
| 878 |
-
},
|
| 879 |
-
"v4": {
|
| 880 |
-
"path": "results/open-reference-gap-v1/nimble-9b/v4/normalized.jsonl",
|
| 881 |
-
"sha256": "6ae263cef9cafdbcd35cd1f1dcc5b1bd920d3e6023d4fc487970286c04ec8d05"
|
| 882 |
-
},
|
| 883 |
-
"v5": {
|
| 884 |
-
"path": "eval/heldout/v5/open-baselines-v1/nimble-9b/core/normalized.jsonl",
|
| 885 |
-
"sha256": "0fc6dae372265ef62ea87ce9dc07b63a77ccd33e3aa48e15754f5ab98c32580a"
|
| 886 |
-
}
|
| 887 |
-
},
|
| 888 |
-
"transfer_rows": {
|
| 889 |
-
"path": "analysis/decision-benchmark-v3-baselines/nimble-9b/ROWS.json",
|
| 890 |
-
"sha256": "f55f29f973b0974adc7fc3fd582e2e5c0414f440e4a094a7ec8e152685e62a28"
|
| 891 |
-
},
|
| 892 |
-
"evidence": [
|
| 893 |
-
{
|
| 894 |
-
"path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/COMPLETE.json",
|
| 895 |
-
"sha256": "efc5a1be536227db99ff9937b84e9a5bc688fc497c0e4a84d8261144cea8a51d"
|
| 896 |
-
},
|
| 897 |
-
{
|
| 898 |
-
"path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/METADATA.json",
|
| 899 |
-
"sha256": "77eaa373ec6a74b483ac9b0cda739ecd8cce95a046266cd7a28488a768b82940"
|
| 900 |
-
},
|
| 901 |
-
{
|
| 902 |
-
"path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/predictions.jsonl",
|
| 903 |
-
"sha256": "ad2212fa691c662c7917b33315439844339252778fdaefcc3080cdf46e00a434"
|
| 904 |
-
},
|
| 905 |
-
{
|
| 906 |
-
"path": "analysis/decision-benchmark-v3-baselines/nimble-9b/REPORT.json",
|
| 907 |
-
"sha256": "b6065189b7522e5d0d2b7b088b673ba26212bee5fda2f0d7c6cd67fc9cc93136"
|
| 908 |
-
},
|
| 909 |
-
{
|
| 910 |
-
"path": "eval/open-reference-comparison-v1/REGISTRY.json",
|
| 911 |
-
"sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
|
| 912 |
-
},
|
| 913 |
-
{
|
| 914 |
-
"path": "analysis/decision-benchmark-v3-baselines/nimble-9b/COMPOSABLE-MANIFEST.json",
|
| 915 |
-
"sha256": "7423a2430984cf8e7eed2c5933ceeb0e2894d252e30e447d1294b6b6b08192bd"
|
| 916 |
-
}
|
| 917 |
-
]
|
| 918 |
}
|
| 919 |
},
|
| 920 |
"task_rows": 54,
|
|
@@ -931,8 +572,7 @@
|
|
| 931 |
"kev-0.8b",
|
| 932 |
"Qwen3.5-2B",
|
| 933 |
"Laya-base",
|
| 934 |
-
"Laya-multilingual"
|
| 935 |
-
"Kai"
|
| 936 |
],
|
| 937 |
"internal_extra_models_not_public_rank": [
|
| 938 |
"llm2jev-2b",
|
|
@@ -1015,5 +655,7 @@
|
|
| 1015 |
}
|
| 1016 |
},
|
| 1017 |
"source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
|
| 1018 |
-
}
|
|
|
|
|
|
|
| 1019 |
}
|
|
|
|
| 1 |
{
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
"statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
|
| 3 |
"manifest_sha256": "3199a40b2719cad9b1e815629cf5ebad923a9a399188c1199a9d27e918bda9e3",
|
| 4 |
"qualified_runtime": {
|
|
|
|
| 219 |
}
|
| 220 |
]
|
| 221 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 222 |
"Laya-base": {
|
| 223 |
"panels": {
|
| 224 |
"v4": {
|
|
|
|
| 556 |
"sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
|
| 557 |
}
|
| 558 |
]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 559 |
}
|
| 560 |
},
|
| 561 |
"task_rows": 54,
|
|
|
|
| 572 |
"kev-0.8b",
|
| 573 |
"Qwen3.5-2B",
|
| 574 |
"Laya-base",
|
| 575 |
+
"Laya-multilingual"
|
|
|
|
| 576 |
],
|
| 577 |
"internal_extra_models_not_public_rank": [
|
| 578 |
"llm2jev-2b",
|
|
|
|
| 655 |
}
|
| 656 |
},
|
| 657 |
"source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
|
| 658 |
+
},
|
| 659 |
+
"frozen_protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
|
| 660 |
+
"presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
|
| 661 |
}
|
release-manifest.json
CHANGED
|
@@ -10,39 +10,41 @@
|
|
| 10 |
"statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
|
| 11 |
"displayed_materials": {
|
| 12 |
"ATTRIBUTIONS.md": "f13d3d134767e99988e721a35500a904ed0fdb72bb48aaf8519b06da3629473d",
|
| 13 |
-
"DIAGNOSTICS.md": "
|
| 14 |
"Dockerfile.runtime": "6c781fdb5ad2abfb8f61c906751ac449d01644c62704cbe28312fd7b7fb57521",
|
| 15 |
-
"EVALUATION.md": "
|
| 16 |
"LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
|
| 17 |
"METHODS.md": "8cc5a8edb9cad84736e19b712a9b9d141883ba7ead5aa3cdd45567180c1db7a4",
|
| 18 |
"QUESTION-SCALING.md": "2a04e4f06132108e022fdc694d9d66c4e30bbd5206149844649138f322e5ca18",
|
| 19 |
"QWEN-LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
|
| 20 |
-
"README.md": "
|
| 21 |
"RUNTIME.md": "91284e3281944e60330258fcbd2617905671da7330f02cac1b0217c987a30095",
|
| 22 |
-
"SENSITIVITY.md": "
|
| 23 |
-
"TASKS.md": "
|
| 24 |
"USAGE.md": "49888ddf4ac8b87d60556a1f44083b4789c9c19d95de460514a26af1f264bc17",
|
| 25 |
-
"WEIGHTING.md": "
|
| 26 |
"assets/architecture.pdf": "482fac3bd270e5155e7d7fe61898c113279214222e4c2124b3e495b04b0fe1ea",
|
| 27 |
"assets/architecture.png": "1e8fbe7655e7c1aadb4b2323290cdb20e726c0a42b7914c0a1d5954a6a029931",
|
| 28 |
"assets/architecture.svg": "76a1f3165173ff07a4126947b9b6ae49da46f64ee2b1d740a3d9a35652fb6bae",
|
| 29 |
-
"assets/decision-matrix.pdf": "
|
| 30 |
-
"assets/decision-matrix.png": "
|
| 31 |
-
"assets/decision-matrix.svg": "
|
| 32 |
"assets/decision-question-scaling.pdf": "d89a843d5b59b1bdd01124ec5e0733a41fa0cf8715ad9adf77cf8e525d749bc2",
|
| 33 |
"assets/decision-question-scaling.png": "fcab4ce4d6ae4876f9110d0231796926faa139a1fbfcd6c3e8ab4f44917be7e7",
|
| 34 |
"assets/decision-question-scaling.svg": "a88c94ace9d3acd6c6171943024d73cdee9e0301d5f4e7287bec488ddcf0b5db",
|
| 35 |
-
"assets/decision-ranking.pdf": "
|
| 36 |
-
"assets/decision-ranking.png": "
|
| 37 |
-
"assets/decision-ranking.svg": "
|
| 38 |
"assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
|
| 39 |
"assets/readout.svg": "701084bf7b0858a3adb247ae446df737b77d4b75004341dfd5f956b933f3d253",
|
| 40 |
-
"metrics/benchmark.json": "
|
| 41 |
-
"metrics/evaluation-provenance.json": "
|
| 42 |
"metrics/question-scaling.json": "859c474c4c1fb571c6041574698f6f47d4982fb56f1f2e68820d8a99b3a41337",
|
| 43 |
"model-card-example.json": "23915dd0232955a2345f53a4e716f43c3816e536a3e5418030db24687644de88",
|
| 44 |
-
"runtime-fla-requirements.lock": "35e1fedff9ca49092a7ada277bbd794abe7e8474e14a678a7e4bcec6406d0eb5"
|
|
|
|
| 45 |
},
|
| 46 |
"weight_identity": "all weights and runtime bound by bundle-manifest.json",
|
| 47 |
-
"shared_diagnostics_materials_sha256": "
|
|
|
|
| 48 |
}
|
|
|
|
| 10 |
"statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
|
| 11 |
"displayed_materials": {
|
| 12 |
"ATTRIBUTIONS.md": "f13d3d134767e99988e721a35500a904ed0fdb72bb48aaf8519b06da3629473d",
|
| 13 |
+
"DIAGNOSTICS.md": "f0e3a5040f9843bf1480f0383d2f01f7e9a601347d6b1545d68bf861d1f64007",
|
| 14 |
"Dockerfile.runtime": "6c781fdb5ad2abfb8f61c906751ac449d01644c62704cbe28312fd7b7fb57521",
|
| 15 |
+
"EVALUATION.md": "0624b6aebb35e86e8456ed734027ae4590fbf24c91775e4e228fe89454780f14",
|
| 16 |
"LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
|
| 17 |
"METHODS.md": "8cc5a8edb9cad84736e19b712a9b9d141883ba7ead5aa3cdd45567180c1db7a4",
|
| 18 |
"QUESTION-SCALING.md": "2a04e4f06132108e022fdc694d9d66c4e30bbd5206149844649138f322e5ca18",
|
| 19 |
"QWEN-LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
|
| 20 |
+
"README.md": "e303076d1f0f3802a170c128d8a10221ddae6169cb0bef52f624811c9f79d5fd",
|
| 21 |
"RUNTIME.md": "91284e3281944e60330258fcbd2617905671da7330f02cac1b0217c987a30095",
|
| 22 |
+
"SENSITIVITY.md": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048",
|
| 23 |
+
"TASKS.md": "ffd8e2da2bed770855e69d54c44479bd66bf27a319684954ad8b55f742706341",
|
| 24 |
"USAGE.md": "49888ddf4ac8b87d60556a1f44083b4789c9c19d95de460514a26af1f264bc17",
|
| 25 |
+
"WEIGHTING.md": "6240f01d42aa4364ca2091f1e2d25619c21db91035348457609d0d2f9a740048",
|
| 26 |
"assets/architecture.pdf": "482fac3bd270e5155e7d7fe61898c113279214222e4c2124b3e495b04b0fe1ea",
|
| 27 |
"assets/architecture.png": "1e8fbe7655e7c1aadb4b2323290cdb20e726c0a42b7914c0a1d5954a6a029931",
|
| 28 |
"assets/architecture.svg": "76a1f3165173ff07a4126947b9b6ae49da46f64ee2b1d740a3d9a35652fb6bae",
|
| 29 |
+
"assets/decision-matrix.pdf": "12f963382e810eb2adb8a9fbbc3b3b7e25e75e4c13d56a04db81419e0b565033",
|
| 30 |
+
"assets/decision-matrix.png": "29c1fc0ad52b53bfc37f76788a9a824d5b767c8ed38c2474b763c1cb941182c4",
|
| 31 |
+
"assets/decision-matrix.svg": "95e90ed75b9d942e212b621a2968be3c5387d0bf637ca681c885062c5de93456",
|
| 32 |
"assets/decision-question-scaling.pdf": "d89a843d5b59b1bdd01124ec5e0733a41fa0cf8715ad9adf77cf8e525d749bc2",
|
| 33 |
"assets/decision-question-scaling.png": "fcab4ce4d6ae4876f9110d0231796926faa139a1fbfcd6c3e8ab4f44917be7e7",
|
| 34 |
"assets/decision-question-scaling.svg": "a88c94ace9d3acd6c6171943024d73cdee9e0301d5f4e7287bec488ddcf0b5db",
|
| 35 |
+
"assets/decision-ranking.pdf": "2c5d77b3a2420531e25a35e74389cd74968840117f90ab9f0f79e3e245101a97",
|
| 36 |
+
"assets/decision-ranking.png": "a201fa8f34771901d05b3162ed33de26f38bdba6dfbcab4579b13de017cce40f",
|
| 37 |
+
"assets/decision-ranking.svg": "39be265fee1f5e6a0b051524fadc33786c0a923a28d771a774ea3c01ace460dc",
|
| 38 |
"assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
|
| 39 |
"assets/readout.svg": "701084bf7b0858a3adb247ae446df737b77d4b75004341dfd5f956b933f3d253",
|
| 40 |
+
"metrics/benchmark.json": "40b1da632cb586a5122fa34f85ea654e2890d6619bbe707b49afe829b4e8d3ce",
|
| 41 |
+
"metrics/evaluation-provenance.json": "cab27fc814933826d4a280ab280877c4683524a5ea67773e4e134a00fcf8befa",
|
| 42 |
"metrics/question-scaling.json": "859c474c4c1fb571c6041574698f6f47d4982fb56f1f2e68820d8a99b3a41337",
|
| 43 |
"model-card-example.json": "23915dd0232955a2345f53a4e716f43c3816e536a3e5418030db24687644de88",
|
| 44 |
+
"runtime-fla-requirements.lock": "35e1fedff9ca49092a7ada277bbd794abe7e8474e14a678a7e4bcec6406d0eb5",
|
| 45 |
+
"MATERIALS.json": "cf00e48f3111ab392aa50fc2b8a0237d27c7aee6a2aaddd696c2af5d8d0cd493"
|
| 46 |
},
|
| 47 |
"weight_identity": "all weights and runtime bound by bundle-manifest.json",
|
| 48 |
+
"shared_diagnostics_materials_sha256": "cf00e48f3111ab392aa50fc2b8a0237d27c7aee6a2aaddd696c2af5d8d0cd493",
|
| 49 |
+
"presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
|
| 50 |
}
|