Focus Lex evaluation on fine-tuned typed-decision models
Browse files- BENCHMARK_COMPARISON.json +0 -0
- BENCHMARK_COMPARISON.md +0 -56
- EVALUATION.md +0 -2
- PACKAGE_MANIFEST.json +7 -16
- README.md +8 -8
BENCHMARK_COMPARISON.json
DELETED
|
The diff for this file is too large to render.
See raw diff
|
|
|
BENCHMARK_COMPARISON.md
DELETED
|
@@ -1,56 +0,0 @@
|
|
| 1 |
-
# Decision family — same-benchmark comparison
|
| 2 |
-
|
| 3 |
-
Fixed public Sol and Nox weights were evaluated once on the original Kai and Lex benchmark samples. Kai, Lex and Laya results reuse their saved predictions; their original aggregations were reproduced before adding the decoder columns. All columns use their released temperature. No model or temperature was fitted on these evaluations.
|
| 4 |
-
|
| 5 |
-
## General decision capability
|
| 6 |
-
|
| 7 |
-
| Text language · decision | Kai | Laya English | Laya Multilingual | Sol | Nox |
|
| 8 |
-
|---|---:|---:|---:|---:|---:|
|
| 9 |
-
| English · Choice | 41.01 | 40.80 | 30.97 | 65.96 | **70.82** |
|
| 10 |
-
| English · Noul (BoolQ)¹ | 71.80 | 72.92 | 62.17 | 74.93 | **85.24** |
|
| 11 |
-
| English · Score | **91.80** | 84.71 | 83.06 | 86.63 | 85.00 |
|
| 12 |
-
| Non-English · Choice | 48.44 | 39.96 | 43.82 | 61.41 | **68.47** |
|
| 13 |
-
| Non-English · Noul | 74.80 | 51.04 | 53.58 | 75.34 | **81.78** |
|
| 14 |
-
| Non-English · Score | **87.47** | 82.08 | 83.41 | 84.29 | 80.10 |
|
| 15 |
-
|
| 16 |
-
Choice is macro accuracy, Noul is balanced accuracy, and Score is 1 − normalized RPS. These are distinct metrics, not a single blended score. The original five-family common window contains 41,720 decisions: Belebele 31,354, XCOPA 5,500, OASST 3,803, Afri 747 and TyDi 316. The English Noul row separately uses 2,303 BoolQ questions. All 44,023 inputs fit both decoder models without truncation. English and non-English aggregation, class balancing and source weights are unchanged.
|
| 17 |
-
|
| 18 |
-
The original common window is Kai/Laya Multilingual ≤1,024 complete tokens and Laya English ≤512. Sol/Nox use their released 16,384-token capacity on these same complete inputs; larger capacity does not change the sample denominator. This table compares general published weights. These panels were already observed during development and are not presented as a new independent holdout.
|
| 19 |
-
|
| 20 |
-
[Original Kai evaluation](https://gist.github.com/Xunzhuo/bfc87193ac13633019fd73cfaccc615f#file-01-kai-evaluation-md).
|
| 21 |
-
|
| 22 |
-
## Typed decisions
|
| 23 |
-
|
| 24 |
-
| Test accuracy | Lex | Laya Typed Decisions | Sol | Nox |
|
| 25 |
-
|---|---:|---:|---:|---:|
|
| 26 |
-
| **Overall** | **78.15%** | 76.60% | 47.55% | 51.50% |
|
| 27 |
-
| Choice · 600 decisions | **74.00%** | 73.33% | 44.50% | 48.33% |
|
| 28 |
-
| Noul · 600 decisions | 84.67% | **85.67%** | 58.50% | 60.50% |
|
| 29 |
-
| Score · 800 decisions | **76.38%** | 72.25% | 41.63% | 47.13% |
|
| 30 |
-
|
| 31 |
-
The original TEST contains 2,000 decisions across 400 state components and four English workflows. **Lex and Laya Typed Decisions are fine-tuned specialists; Sol and Nox are general published models.** This describes their fixed releases on the same tasks, rather than equal-budget fine-tuning. The TEST remains excluded from subsequent Lex recipe, checkpoint and temperature selection.
|
| 32 |
-
|
| 33 |
-
| Model | Soft NLL ↓ | Mean Brier ↓ | Score RPS ↓ | Score MAE ↓ |
|
| 34 |
-
|---|---:|---:|---:|---:|
|
| 35 |
-
| Lex | 0.961902 | 0.030858 | 0.022994 | 0.251408 |
|
| 36 |
-
| Laya Typed Decisions | 0.884390 | 0.018452 | 0.015602 | 0.242408 |
|
| 37 |
-
| Sol | 1.357052 | 0.094882 | 0.081148 | 0.601998 |
|
| 38 |
-
| Nox | 2.375962 | 0.136159 | 0.117172 | 0.676995 |
|
| 39 |
-
|
| 40 |
-
Soft metrics measure agreement with the original normalized soft labels. Score MAE uses the original numerical rubric values. These do not establish calibration for an arbitrary application.
|
| 41 |
-
|
| 42 |
-
| Paired accuracy difference | Points | 95% interval |
|
| 43 |
-
|---|---:|---:|
|
| 44 |
-
| Nox minus Lex | -26.65 | [-29.35, -24.05] |
|
| 45 |
-
| Nox minus Laya Typed Decisions | -25.10 | [-27.80, -22.45] |
|
| 46 |
-
| Sol minus Lex | -30.60 | [-33.20, -28.05] |
|
| 47 |
-
| Sol minus Laya Typed Decisions | -29.05 | [-31.70, -26.45] |
|
| 48 |
-
|
| 49 |
-
Intervals resample whole state components using the original 2,000-replicate stratified bootstrap. No new confidence intervals are claimed for the six general panels. [Original Lex evaluation](https://gist.github.com/Xunzhuo/76b59cb158ce5069a53b946c6f0ee656#file-01-lex-evaluation-md).
|
| 50 |
-
|
| 51 |
-
## Fixed releases and execution
|
| 52 |
-
|
| 53 |
-
- [Nox · c9bad4f4](https://huggingface.co/llm-semantic-router/Decision-1.0-Nox/tree/c9bad4f4b8bafa56522f1ce1a90dcbe1e44df0a0)
|
| 54 |
-
- [Sol · d29e9a71](https://huggingface.co/llm-semantic-router/Decision-1.0-Sol/tree/d29e9a71e3e7feeac667140a7180e9cfdcabc0b2)
|
| 55 |
-
|
| 56 |
-
Each decoder completed 46,023 decisions in 5,753 AMD B8 forwards, with zero unsupported inputs and no parameter updates. Sol/Nox retain their public BF16 backbone and FP32 head; original encoder predictions use their existing FP32 runtime. This is an accuracy comparison, not a serving-latency measurement. Raw and shipped-temperature scores are retained in [structured results](BENCHMARK_COMPARISON.json).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
EVALUATION.md
CHANGED
|
@@ -1,7 +1,5 @@
|
|
| 1 |
# Official typed-decisions TEST
|
| 2 |
|
| 3 |
-
[Expanded comparison with fixed published Sol and Nox](BENCHMARK_COMPARISON.md). Original encoder evaluation follows unchanged.
|
| 4 |
-
|
| 5 |
The fixed final Lex checkpoint and the official `laya-typed-decisions` specialist were each evaluated once on the original 2,000 TEST decisions, from 400 state components. Both models support all 2,000 complete inputs within 1,024 tokens, their native windows and their untruncated official-default support. No TRAIN state component or full input overlaps this TEST support. The test contains four English workflows and 20 question schemas. This is not evidence of multilingual or unseen-schema generalization.
|
| 6 |
|
| 7 |
| Metric | Lex | Official specialist (shipped temperature) |
|
|
|
|
| 1 |
# Official typed-decisions TEST
|
| 2 |
|
|
|
|
|
|
|
| 3 |
The fixed final Lex checkpoint and the official `laya-typed-decisions` specialist were each evaluated once on the original 2,000 TEST decisions, from 400 state components. Both models support all 2,000 complete inputs within 1,024 tokens, their native windows and their untruncated official-default support. No TRAIN state component or full input overlaps this TEST support. The test contains four English workflows and 20 question schemas. This is not evidence of multilingual or unseen-schema generalization.
|
| 4 |
|
| 5 |
| Metric | Lex | Official specialist (shipped temperature) |
|
PACKAGE_MANIFEST.json
CHANGED
|
@@ -13,8 +13,8 @@
|
|
| 13 |
"sha256": "54a5895871af33845ee939f2e8f37b7f821cee8e89ff25ec1fd78e9cf83e9c63"
|
| 14 |
},
|
| 15 |
"EVALUATION.md": {
|
| 16 |
-
"bytes":
|
| 17 |
-
"sha256": "
|
| 18 |
},
|
| 19 |
"FINETUNING.md": {
|
| 20 |
"bytes": 3629,
|
|
@@ -61,8 +61,8 @@
|
|
| 61 |
"sha256": "2e5936d3ea196e00968cc37b7662a6afbe8cfd6fd11d04b9fecba554b1e1e37b"
|
| 62 |
},
|
| 63 |
"README.md": {
|
| 64 |
-
"bytes":
|
| 65 |
-
"sha256": "
|
| 66 |
},
|
| 67 |
"TECHNICAL_VALIDATION.json": {
|
| 68 |
"bytes": 1547,
|
|
@@ -363,14 +363,6 @@
|
|
| 363 |
"LICENSES/Upstream-MIT.txt": {
|
| 364 |
"bytes": 1191,
|
| 365 |
"sha256": "3ec44d2f046b27986e28e3b4705110a33b139e6392ef3868ff594a0453311038"
|
| 366 |
-
},
|
| 367 |
-
"BENCHMARK_COMPARISON.md": {
|
| 368 |
-
"bytes": 4485,
|
| 369 |
-
"sha256": "ba0c0eb7f7c65b584ceee6ba2d08068a6e188e904c32856df136202dce98763b"
|
| 370 |
-
},
|
| 371 |
-
"BENCHMARK_COMPARISON.json": {
|
| 372 |
-
"bytes": 609635,
|
| 373 |
-
"sha256": "45990555c872f93ef866f59653019b6876d7ca44b9d430f44d0c084d46a5ceb1"
|
| 374 |
}
|
| 375 |
},
|
| 376 |
"language_scope": "English typed-decisions specialist",
|
|
@@ -385,12 +377,11 @@
|
|
| 385 |
"subject_manifest_sha256": "f288d873999832a3f37c6a7c4268c2ab309691e621794dbf7acab891acbbb7e6",
|
| 386 |
"technical_validation": "TECHNICAL_VALIDATION.json",
|
| 387 |
"training_provenance": "TRAINING_PROVENANCE.json",
|
| 388 |
-
"file_count_excluding_this_manifest":
|
| 389 |
-
"bytes_excluding_this_manifest":
|
| 390 |
"publication": {
|
| 391 |
"download_access": "public_ungated",
|
| 392 |
"new_contribution_license": "Apache-2.0",
|
| 393 |
"retained_third_party_terms": true
|
| 394 |
-
}
|
| 395 |
-
"public_non_native_bytes_excluding_this_manifest": 5696013
|
| 396 |
}
|
|
|
|
| 13 |
"sha256": "54a5895871af33845ee939f2e8f37b7f821cee8e89ff25ec1fd78e9cf83e9c63"
|
| 14 |
},
|
| 15 |
"EVALUATION.md": {
|
| 16 |
+
"bytes": 2813,
|
| 17 |
+
"sha256": "00862d0a49c69c406b7dbfa890f95f09a34680568309d1cc0571733abe7b933c"
|
| 18 |
},
|
| 19 |
"FINETUNING.md": {
|
| 20 |
"bytes": 3629,
|
|
|
|
| 61 |
"sha256": "2e5936d3ea196e00968cc37b7662a6afbe8cfd6fd11d04b9fecba554b1e1e37b"
|
| 62 |
},
|
| 63 |
"README.md": {
|
| 64 |
+
"bytes": 3816,
|
| 65 |
+
"sha256": "43406a86a3c38ccda748d6c4f96e7be18af9dca8aaf3229c283154b102320b70"
|
| 66 |
},
|
| 67 |
"TECHNICAL_VALIDATION.json": {
|
| 68 |
"bytes": 1547,
|
|
|
|
| 363 |
"LICENSES/Upstream-MIT.txt": {
|
| 364 |
"bytes": 1191,
|
| 365 |
"sha256": "3ec44d2f046b27986e28e3b4705110a33b139e6392ef3868ff594a0453311038"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 366 |
}
|
| 367 |
},
|
| 368 |
"language_scope": "English typed-decisions specialist",
|
|
|
|
| 377 |
"subject_manifest_sha256": "f288d873999832a3f37c6a7c4268c2ab309691e621794dbf7acab891acbbb7e6",
|
| 378 |
"technical_validation": "TECHNICAL_VALIDATION.json",
|
| 379 |
"training_provenance": "TRAINING_PROVENANCE.json",
|
| 380 |
+
"file_count_excluding_this_manifest": 91,
|
| 381 |
+
"bytes_excluding_this_manifest": 2327340742,
|
| 382 |
"publication": {
|
| 383 |
"download_access": "public_ungated",
|
| 384 |
"new_contribution_license": "Apache-2.0",
|
| 385 |
"retained_third_party_terms": true
|
| 386 |
+
}
|
|
|
|
| 387 |
}
|
README.md
CHANGED
|
@@ -29,16 +29,16 @@ Apache 2.0 · No access request · [Try Decision Studio](https://llm-semantic-ro
|
|
| 29 |
|
| 30 |
**78.15% accuracy — 1.55 points above Laya Typed Decisions.**
|
| 31 |
|
| 32 |
-
|
| 33 |
|
| 34 |
-
| Test accuracy | Lex | Laya Typed Decisions |
|
| 35 |
-
|---|---:|---:|
|
| 36 |
-
| **Overall** | **78.15%** | 76.60% |
|
| 37 |
-
| Choice · 600 decisions | **74.00%** | 73.33% |
|
| 38 |
-
| Noul · 600 decisions | 84.67% | **85.67%** |
|
| 39 |
-
| Score · 800 decisions | **76.38%** | 72.25% |
|
| 40 |
|
| 41 |
-
|
| 42 |
|
| 43 |
## Built for operational decisions
|
| 44 |
|
|
|
|
| 29 |
|
| 30 |
**78.15% accuracy — 1.55 points above Laya Typed Decisions.**
|
| 31 |
|
| 32 |
+
Both models were evaluated on the original 2,000-decision typed-decisions test split, with complete inputs and the official fine-tuned Laya checkpoint. Lex answers **31 more decisions correctly**.
|
| 33 |
|
| 34 |
+
| Test accuracy | Lex | Laya Typed Decisions |
|
| 35 |
+
|---|---:|---:|
|
| 36 |
+
| **Overall** | **78.15%** | 76.60% |
|
| 37 |
+
| Choice · 600 decisions | **74.00%** | 73.33% |
|
| 38 |
+
| Noul · 600 decisions | 84.67% | **85.67%** |
|
| 39 |
+
| Score · 800 decisions | **76.38%** | 72.25% |
|
| 40 |
|
| 41 |
+
Observed accuracy gain; the paired 95% interval is −0.20 to +3.15 points. Laya leads on Noul and probability-quality metrics. [Full evaluation](https://gist.github.com/Xunzhuo/76b59cb158ce5069a53b946c6f0ee656#file-01-lex-evaluation-md).
|
| 42 |
|
| 43 |
## Built for operational decisions
|
| 44 |
|