Publish qualified Nox Choice semantics and latest Decision comparison
Browse files- DIAGNOSTICS.md +8 -8
- EVALUATION.md +3 -3
- MATERIALS.json +22 -26
- NULL_DESCRIPTION_RENDERING.json +7 -0
- README.md +5 -3
- RUNTIME-RELEASE.json +9 -28
- SENSITIVITY.md +3 -3
- TASKS.md +64 -64
- USAGE.md +3 -1
- WEIGHTING.md +17 -0
- assets/decision-matrix.pdf +0 -0
- assets/decision-matrix.png +2 -2
- assets/decision-matrix.svg +108 -108
- assets/decision-ranking.pdf +0 -0
- assets/decision-ranking.png +2 -2
- assets/decision-ranking.svg +23 -23
- bundle-manifest.json +10 -5
- code/decision_api.py +1 -1
- metrics/benchmark.json +518 -1159
- metrics/evaluation-provenance.json +272 -257
- model-card-example.json +1 -1
- release-manifest.json +54 -44
DIAGNOSTICS.md
CHANGED
|
@@ -6,9 +6,8 @@ These axes remain separate from headline accuracy. Probability metrics use the s
|
|
| 6 |
|
| 7 |
| Model | Valid / requested | Brier ↓ | NLL ↓ | ECE % ↓ | Coverage at ≤5% error % ↑ | AURC ↓ |
|
| 8 |
|---|---:|---:|---:|---:|---:|---:|
|
| 9 |
-
| Jev | 1046/1046 | 0.1912 | 0.5743 | 3.35 | 76.96 | 0.0346 |
|
| 10 |
| Lux | 1046/1046 | 0.3058 | 0.6201 | 5.13 | 57.36 | 0.0639 |
|
| 11 |
-
| Nox | 1046/1046 | 0.
|
| 12 |
| Kev-9B | 1046/1046 | 0.2948 | 0.6028 | 4.45 | 58.03 | 0.0586 |
|
| 13 |
| Kev-4B | 1046/1046 | 0.3145 | 0.6434 | 2.61 | 56.79 | 0.0659 |
|
| 14 |
| Qwen3.5-9B | 1046/1046 | 0.3673 | 0.7597 | 9.75 | 47.42 | 0.0872 |
|
|
@@ -19,6 +18,7 @@ These axes remain separate from headline accuracy. Probability metrics use the s
|
|
| 19 |
| Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
|
| 20 |
| Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
|
| 21 |
| Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
|
|
|
|
| 22 |
|
| 23 |
Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
|
| 24 |
|
|
@@ -26,9 +26,8 @@ Coverage at an error threshold keeps whole confidence-tie groups together. These
|
|
| 26 |
|
| 27 |
| Model | Valid / requested pairs | Both correct % ↑ | Semantic flip % ↓ | Mean half-L1 ↓ |
|
| 28 |
|---|---:|---:|---:|---:|
|
| 29 |
-
| Jev | 36/36 | 86.11 | 0.00 | 0.0208 |
|
| 30 |
| Lux | 36/36 | 77.78 | 8.33 | 0.0923 |
|
| 31 |
-
| Nox | 36/36 |
|
| 32 |
| Kev-9B | 36/36 | 80.56 | 2.78 | 0.0615 |
|
| 33 |
| Kev-4B | 36/36 | 77.78 | 5.56 | 0.0667 |
|
| 34 |
| Qwen3.5-9B | 36/36 | 75.00 | 11.11 | 0.1157 |
|
|
@@ -39,6 +38,7 @@ Coverage at an error threshold keeps whole confidence-tie groups together. These
|
|
| 39 |
| Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
|
| 40 |
| Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
|
| 41 |
| Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
|
|
|
|
| 42 |
|
| 43 |
The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
|
| 44 |
|
|
@@ -46,7 +46,6 @@ The 36 paired permutations test the same semantics under changed option order. S
|
|
| 46 |
|
| 47 |
| Model | Valid / requested | Intact/control accuracy % ↑ | Mean max P % ↓ | P≥0.9 share % ↓ | Normalized entropy ↑ | Paired confidence drop pp ↑ |
|
| 48 |
|---|---:|---:|---:|---:|---:|---:|
|
| 49 |
-
| Jev | 110/110 | 93.64 | 62.04 | 16.36 | 0.7657 | 30.30 |
|
| 50 |
| Lux | 110/110 | 84.55 | 68.81 | 19.09 | 0.6785 | 19.59 |
|
| 51 |
| Nox | 110/110 | 72.73 | 78.65 | 27.27 | 0.5072 | 12.64 |
|
| 52 |
| Kev-9B | 110/110 | 91.82 | 39.61 | 0.00 | 0.9981 | 53.19 |
|
|
@@ -59,6 +58,7 @@ The 36 paired permutations test the same semantics under changed option order. S
|
|
| 59 |
| Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
|
| 60 |
| Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
|
| 61 |
| Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
|
|
|
|
| 62 |
|
| 63 |
The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
|
| 64 |
|
|
@@ -66,7 +66,6 @@ The 110 unknowable examples have no scored true class and are excluded from accu
|
|
| 66 |
|
| 67 |
| Model | Original probability rows | Transfer probability rows | Transfer truncated questions |
|
| 68 |
|---|---:|---:|---:|
|
| 69 |
-
| Jev | 2720/2720 | 1264/1264 | — / not observable |
|
| 70 |
| Lux | 2720/2720 | 1264/1264 | 0 |
|
| 71 |
| Nox | 2720/2720 | 1264/1264 | 0 |
|
| 72 |
| Kev-9B | 2720/2720 | 1264/1264 | 0 |
|
|
@@ -79,6 +78,7 @@ The 110 unknowable examples have no scored true class and are excluded from accu
|
|
| 79 |
| Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
|
| 80 |
| Laya · English | 2720/2720 | 1264/1264 | 34 |
|
| 81 |
| Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
|
|
|
|
| 82 |
|
| 83 |
Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation. Server-side truncation for Jev cannot be observed.
|
| 84 |
|
|
@@ -86,9 +86,8 @@ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and orde
|
|
| 86 |
|
| 87 |
| Model | Overall % | 95% component-bootstrap interval |
|
| 88 |
|---|---:|---:|
|
| 89 |
-
| Jev | 81.05 | 79.70–82.35 |
|
| 90 |
| Lux | 76.72 | 75.35–78.07 |
|
| 91 |
-
| Nox |
|
| 92 |
| Kev-9B | 71.89 | 70.42–73.35 |
|
| 93 |
| Kev-4B | 70.09 | 68.45–71.63 |
|
| 94 |
| Qwen3.5-9B | 69.73 | 68.27–71.20 |
|
|
@@ -99,5 +98,6 @@ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and orde
|
|
| 99 |
| Qwen3.5-2B | 57.24 | 55.74–58.76 |
|
| 100 |
| Laya · English | 51.03 | 49.43–52.68 |
|
| 101 |
| Laya · Multilingual | 47.19 | 45.58–48.82 |
|
|
|
|
| 102 |
|
| 103 |
Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
|
|
|
|
| 6 |
|
| 7 |
| Model | Valid / requested | Brier ↓ | NLL ↓ | ECE % ↓ | Coverage at ≤5% error % ↑ | AURC ↓ |
|
| 8 |
|---|---:|---:|---:|---:|---:|---:|
|
|
|
|
| 9 |
| Lux | 1046/1046 | 0.3058 | 0.6201 | 5.13 | 57.36 | 0.0639 |
|
| 10 |
+
| Nox | 1046/1046 | 0.4169 | 0.9079 | 11.80 | 36.42 | 0.1109 |
|
| 11 |
| Kev-9B | 1046/1046 | 0.2948 | 0.6028 | 4.45 | 58.03 | 0.0586 |
|
| 12 |
| Kev-4B | 1046/1046 | 0.3145 | 0.6434 | 2.61 | 56.79 | 0.0659 |
|
| 13 |
| Qwen3.5-9B | 1046/1046 | 0.3673 | 0.7597 | 9.75 | 47.42 | 0.0872 |
|
|
|
|
| 18 |
| Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
|
| 19 |
| Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
|
| 20 |
| Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
|
| 21 |
+
| Jev | 1046/1046 | 0.1912 | 0.5743 | 3.35 | 76.96 | 0.0346 |
|
| 22 |
|
| 23 |
Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
|
| 24 |
|
|
|
|
| 26 |
|
| 27 |
| Model | Valid / requested pairs | Both correct % ↑ | Semantic flip % ↓ | Mean half-L1 ↓ |
|
| 28 |
|---|---:|---:|---:|---:|
|
|
|
|
| 29 |
| Lux | 36/36 | 77.78 | 8.33 | 0.0923 |
|
| 30 |
+
| Nox | 36/36 | 63.89 | 16.67 | 0.0795 |
|
| 31 |
| Kev-9B | 36/36 | 80.56 | 2.78 | 0.0615 |
|
| 32 |
| Kev-4B | 36/36 | 77.78 | 5.56 | 0.0667 |
|
| 33 |
| Qwen3.5-9B | 36/36 | 75.00 | 11.11 | 0.1157 |
|
|
|
|
| 38 |
| Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
|
| 39 |
| Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
|
| 40 |
| Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
|
| 41 |
+
| Jev | 36/36 | 86.11 | 0.00 | 0.0208 |
|
| 42 |
|
| 43 |
The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
|
| 44 |
|
|
|
|
| 46 |
|
| 47 |
| Model | Valid / requested | Intact/control accuracy % ↑ | Mean max P % ↓ | P≥0.9 share % ↓ | Normalized entropy ↑ | Paired confidence drop pp ↑ |
|
| 48 |
|---|---:|---:|---:|---:|---:|---:|
|
|
|
|
| 49 |
| Lux | 110/110 | 84.55 | 68.81 | 19.09 | 0.6785 | 19.59 |
|
| 50 |
| Nox | 110/110 | 72.73 | 78.65 | 27.27 | 0.5072 | 12.64 |
|
| 51 |
| Kev-9B | 110/110 | 91.82 | 39.61 | 0.00 | 0.9981 | 53.19 |
|
|
|
|
| 58 |
| Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
|
| 59 |
| Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
|
| 60 |
| Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
|
| 61 |
+
| Jev | 110/110 | 93.64 | 62.04 | 16.36 | 0.7657 | 30.30 |
|
| 62 |
|
| 63 |
The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
|
| 64 |
|
|
|
|
| 66 |
|
| 67 |
| Model | Original probability rows | Transfer probability rows | Transfer truncated questions |
|
| 68 |
|---|---:|---:|---:|
|
|
|
|
| 69 |
| Lux | 2720/2720 | 1264/1264 | 0 |
|
| 70 |
| Nox | 2720/2720 | 1264/1264 | 0 |
|
| 71 |
| Kev-9B | 2720/2720 | 1264/1264 | 0 |
|
|
|
|
| 78 |
| Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
|
| 79 |
| Laya · English | 2720/2720 | 1264/1264 | 34 |
|
| 80 |
| Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
|
| 81 |
+
| Jev | 2720/2720 | 1264/1264 | — / not observable |
|
| 82 |
|
| 83 |
Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation. Server-side truncation for Jev cannot be observed.
|
| 84 |
|
|
|
|
| 86 |
|
| 87 |
| Model | Overall % | 95% component-bootstrap interval |
|
| 88 |
|---|---:|---:|
|
|
|
|
| 89 |
| Lux | 76.72 | 75.35–78.07 |
|
| 90 |
+
| Nox | 73.09 | 71.57–74.56 |
|
| 91 |
| Kev-9B | 71.89 | 70.42–73.35 |
|
| 92 |
| Kev-4B | 70.09 | 68.45–71.63 |
|
| 93 |
| Qwen3.5-9B | 69.73 | 68.27–71.20 |
|
|
|
|
| 98 |
| Qwen3.5-2B | 57.24 | 55.74–58.76 |
|
| 99 |
| Laya · English | 51.03 | 49.43–52.68 |
|
| 100 |
| Laya · Multilingual | 47.19 | 45.58–48.82 |
|
| 101 |
+
| Jev | 81.05 | 79.70–82.35 |
|
| 102 |
|
| 103 |
Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
|
EVALUATION.md
CHANGED
|
@@ -4,9 +4,8 @@ The comparison covers **3,766 scored decisions across 54 tasks** and all 13 disp
|
|
| 4 |
|
| 5 |
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|
| 6 |
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 7 |
-
|
|
| 8 |
| Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
|
| 9 |
-
| Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** |
|
| 10 |
| Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
|
| 11 |
| Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
|
| 12 |
| Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
|
|
@@ -17,8 +16,9 @@ The comparison covers **3,766 scored decisions across 54 tasks** and all 13 disp
|
|
| 17 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
|
| 18 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
|
| 19 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
|
|
|
|
| 20 |
|
| 21 |
-
Accuracy (%).
|
| 22 |
|
| 23 |
## Scope and weighting
|
| 24 |
|
|
|
|
| 4 |
|
| 5 |
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|
| 6 |
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 7 |
+
| Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 69.60 | **73.09** |
|
| 8 |
| Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
|
|
|
|
| 9 |
| Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
|
| 10 |
| Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
|
| 11 |
| Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
|
|
|
|
| 16 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
|
| 17 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
|
| 18 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
|
| 19 |
+
| Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
|
| 20 |
|
| 21 |
+
Accuracy (%). The current model is first, other open models follow by overall score, and Jev is the frontier reference at the end. Bold marks Decision-family cells strictly above every external open or untuned reference in that metric; Jev and the other Decision models are excluded from that threshold.
|
| 22 |
|
| 23 |
## Scope and weighting
|
| 24 |
|
MATERIALS.json
CHANGED
|
@@ -1,30 +1,26 @@
|
|
| 1 |
{
|
| 2 |
-
"scope": "
|
| 3 |
-
"
|
| 4 |
-
"task_rows": 54,
|
| 5 |
-
"presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f",
|
| 6 |
-
"weights_runtime_temperature_unchanged": true,
|
| 7 |
"files": {
|
| 8 |
-
"
|
| 9 |
-
"EVALUATION.md": "
|
| 10 |
-
"
|
| 11 |
-
"
|
| 12 |
-
"SENSITIVITY.md": "
|
| 13 |
-
"
|
| 14 |
-
"
|
| 15 |
-
"
|
| 16 |
-
"assets/decision-
|
| 17 |
-
"assets/decision-
|
| 18 |
-
"assets/decision-
|
| 19 |
-
"assets/decision-
|
| 20 |
-
"assets/decision-
|
| 21 |
-
"assets/decision-
|
| 22 |
-
"
|
| 23 |
-
"
|
| 24 |
-
"
|
| 25 |
-
"
|
| 26 |
-
"
|
| 27 |
-
"
|
| 28 |
-
"metrics/question-scaling.json": "9407f7ebf86581cdea21e704cdbebb9ced0a485eb980925f48233ebea0f830dc"
|
| 29 |
}
|
| 30 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"scope": "Latest13model comparison,54tasks and separate diagnostics.",
|
| 3 |
+
"statistics_sha256": "b319a993ac569ea1fa8da027ccb2195868557e4efe232a5d06091db2884086be",
|
|
|
|
|
|
|
|
|
|
| 4 |
"files": {
|
| 5 |
+
"README.md": "41dfa047d33bdc3aa9feff55928972533ba803a30b332ca380f042268fa1e848",
|
| 6 |
+
"EVALUATION.md": "4159b7939a107e91b13e47345a192eea489f8bd870943e4b397e094b33bba905",
|
| 7 |
+
"TASKS.md": "2fb04bc9c17d805e2abd5e90520df25e79d85aa5296cd42074d4b156f54bf862",
|
| 8 |
+
"DIAGNOSTICS.md": "0b290793565cc5007e9ee7bc79625812c73202c7d8bfce79d8816a0e7a93a965",
|
| 9 |
+
"SENSITIVITY.md": "c3ea31420db4f2e5f17cc7197197a3ddd18abdad773ffb5bf1171b67136c3bc2",
|
| 10 |
+
"WEIGHTING.md": "c3ea31420db4f2e5f17cc7197197a3ddd18abdad773ffb5bf1171b67136c3bc2",
|
| 11 |
+
"metrics/benchmark.json": "178e2e5f58da45f9fbf378ca82b49d7878fc8e9c93eee195a5ef1e7ba0f4fc6b",
|
| 12 |
+
"metrics/evaluation-provenance.json": "a3718fcee42a8cfe0144b01e04346c36a09ac7958154e4561327753d87d5fe8c",
|
| 13 |
+
"assets/decision-ranking.svg": "ab7ca66a9202587dee593714e46be760cb12ac91b6459074bcbbe11c59107d9d",
|
| 14 |
+
"assets/decision-ranking.pdf": "a28d4b0e8649cb7bf2db452a11d93675631af795561726aa63432b05b5888049",
|
| 15 |
+
"assets/decision-ranking.png": "d252593501e8faeb9c5c8cc0a9ae9a96591e7aa8696cf5394cdfec9c0ddac3b7",
|
| 16 |
+
"assets/decision-matrix.svg": "d99102fafe842f2a6bdd4d455f915833d5e0e09e3a22e5211fbf86e5647247ab",
|
| 17 |
+
"assets/decision-matrix.pdf": "5e127c0cd1a6f83b772a3e66ac2e1f4ec19913ac4938371c6f190a23c50a29fe",
|
| 18 |
+
"assets/decision-matrix.png": "c775737aedeecadb0824092d37014a8e2cffa90af0dafcf10d584e0db823ecf8",
|
| 19 |
+
"code/decision_api.py": "6716bdabca3cd2f1aa447d42a6cf53ee62d7ea97984a01f16b4c09fac8aaf17e",
|
| 20 |
+
"bundle-manifest.json": "7d9b06bc25a75b1f131aabd280df0ef2777bf69db98574a300067ace40c61c9c",
|
| 21 |
+
"NULL_DESCRIPTION_RENDERING.json": "3b531cab60cba35648ab71fe4be7fd1af619f02ee337e3beed6296fdfc8b9ec8",
|
| 22 |
+
"model-card-example.json": "1c4fc87d543722e2c3b839cbebd30c5d1370fcaf2d897f97de13079a668ac93e",
|
| 23 |
+
"USAGE.md": "b4c2f44ce120149c7cdee3749283958f34ba9caf4b6492442818c1b6a748c0d3",
|
| 24 |
+
"RUNTIME-RELEASE.json": "325d7e037acab5f61dab4d56d817e7d7e2e873d9b5f0d0df587afd6c33364a46"
|
|
|
|
| 25 |
}
|
| 26 |
}
|
NULL_DESCRIPTION_RENDERING.json
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"rule": "For Choice only, replace a null description with its original key text before tokenization.",
|
| 3 |
+
"preserved": "Non-null descriptions including empty string; keys/order; Noul and Score; original request; weights/tokenizer/temperature/numeric runtime.",
|
| 4 |
+
"source_bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
|
| 5 |
+
"qualification": "Candidate only; full fixed benchmark and independent public loading proof required.",
|
| 6 |
+
"not_a_weight_training_update": true
|
| 7 |
+
}
|
README.md
CHANGED
|
@@ -32,13 +32,12 @@ tags:
|
|
| 32 |
|
| 33 |
## Measured capability
|
| 34 |
|
| 35 |
-
**
|
| 36 |
|
| 37 |
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|
| 38 |
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 39 |
-
|
|
| 40 |
| Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
|
| 41 |
-
| Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** |
|
| 42 |
| Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
|
| 43 |
| Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
|
| 44 |
| Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
|
|
@@ -49,6 +48,7 @@ tags:
|
|
| 49 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
|
| 50 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
|
| 51 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
|
|
|
|
| 52 |
|
| 53 |
Accuracy (%). Overall weights: Decisions **30%**, Composition **25%**, Reading **15%**, Inference **15%**, Transfer **15%**. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.
|
| 54 |
|
|
@@ -78,6 +78,8 @@ print(model.decide(**REQUEST)["answers"])
|
|
| 78 |
|
| 79 |
[Tested SystemOne request and output](model-card-example.json) · [Install and typed API guide](USAGE.md)
|
| 80 |
|
|
|
|
|
|
|
| 81 |
The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.
|
| 82 |
|
| 83 |
## Architecture
|
|
|
|
| 32 |
|
| 33 |
## Measured capability
|
| 34 |
|
| 35 |
+
**73.09% overall accuracy** across 3,766 scored decisions and 54 tasks. Nox leads Kev-4B by **2.99 percentage points** on this decision-focused comparison; reading and transfer remain opportunities to improve.
|
| 36 |
|
| 37 |
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|
| 38 |
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 39 |
+
| Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 69.60 | **73.09** |
|
| 40 |
| Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
|
|
|
|
| 41 |
| Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
|
| 42 |
| Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
|
| 43 |
| Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
|
|
|
|
| 48 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
|
| 49 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
|
| 50 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
|
| 51 |
+
| Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
|
| 52 |
|
| 53 |
Accuracy (%). Overall weights: Decisions **30%**, Composition **25%**, Reading **15%**, Inference **15%**, Transfer **15%**. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.
|
| 54 |
|
|
|
|
| 78 |
|
| 79 |
[Tested SystemOne request and output](model-card-example.json) · [Install and typed API guide](USAGE.md)
|
| 80 |
|
| 81 |
+
Choice candidates with a null description use their ID text, which may increase input tokens.
|
| 82 |
+
|
| 83 |
The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.
|
| 84 |
|
| 85 |
## Architecture
|
RUNTIME-RELEASE.json
CHANGED
|
@@ -1,30 +1,11 @@
|
|
| 1 |
{
|
| 2 |
-
"
|
| 3 |
-
"
|
| 4 |
-
"
|
| 5 |
-
"
|
| 6 |
-
"
|
| 7 |
-
"
|
| 8 |
-
"
|
| 9 |
-
"
|
| 10 |
-
|
| 11 |
-
],
|
| 12 |
-
"weights_tokenizer_prompts_calibration_profile_unchanged": true,
|
| 13 |
-
"default_public_parity": {
|
| 14 |
-
"requests": 2856,
|
| 15 |
-
"answers": 3160,
|
| 16 |
-
"fixtures": 58,
|
| 17 |
-
"raw_logits_probabilities_typed_outputs_exact": true,
|
| 18 |
-
"receipt_sha256": "946dba5c31074b63efa3223ede17857e5f9d1b7b3b44de08cd9add9ceece878c"
|
| 19 |
-
},
|
| 20 |
-
"timing": {
|
| 21 |
-
"receipt_sha256": "501f25efd1813b566d77e98278229440b7a3b4a2358def1462c0171a33e9f3be",
|
| 22 |
-
"shape": "distinct fixed-length questions; Q=32; 499 tokens per question",
|
| 23 |
-
"reference_ms": 419.124,
|
| 24 |
-
"candidate_ms": 414.199,
|
| 25 |
-
"same_measured_api_sha256": "273f6f10f22d5a68b8db34cfcbd35407fb43d8030f6d7cd188bcf118dd90a152",
|
| 26 |
-
"no_shared_state_neural_cache_claim": true
|
| 27 |
-
},
|
| 28 |
-
"capability_scores_unchanged": true,
|
| 29 |
-
"downloaded_package_offline_proof_required_before_promotion": true
|
| 30 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"release": "v1.3.2",
|
| 3 |
+
"kind": "SystemOne Choice null-description semantics",
|
| 4 |
+
"rule": "A null Choice description uses the original candidate key as its description.",
|
| 5 |
+
"weights_tokenizer_temperature_unchanged": true,
|
| 6 |
+
"explicit_descriptions_and_Noul_Score_unchanged": true,
|
| 7 |
+
"bundle_manifest_sha256": "7d9b06bc25a75b1f131aabd280df0ef2777bf69db98574a300067ace40c61c9c",
|
| 8 |
+
"qualified_statistics_sha256": "b319a993ac569ea1fa8da027ccb2195868557e4efe232a5d06091db2884086be",
|
| 9 |
+
"exact_public_contract_proof_sha256": "ebf27c7aa2bb4a12ba2cbb27e4096e46cb43ac4a20b9d3e1bc34f62777a1e63a",
|
| 10 |
+
"runtime_note": "This is an API rendering improvement, not a newly trained checkpoint."
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
}
|
SENSITIVITY.md
CHANGED
|
@@ -1,10 +1,9 @@
|
|
| 1 |
-
Product weights were
|
| 2 |
|
| 3 |
| Model | Current 30/25/15/15/15 | Earlier 25/25/15/15/20 | Original four-panel mean |
|
| 4 |
|---|---:|---:|---:|
|
| 5 |
-
| Jev | 81.05 | 81.45 | 82.45 |
|
| 6 |
| Lux | 76.72 | 76.43 | 79.06 |
|
| 7 |
-
| Nox |
|
| 8 |
| Kev-9B | 71.89 | 72.01 | 73.19 |
|
| 9 |
| Kev-4B | 70.09 | 70.30 | 71.73 |
|
| 10 |
| Qwen3.5-9B | 69.73 | 69.70 | 71.99 |
|
|
@@ -15,3 +14,4 @@ Product weights were changed after earlier results were observed. These are iden
|
|
| 15 |
| Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
|
| 16 |
| Laya · English | 51.03 | 50.85 | 51.76 |
|
| 17 |
| Laya · Multilingual | 47.19 | 47.18 | 48.56 |
|
|
|
|
|
|
| 1 |
+
The same latest model predictions are shown under three weighting schemes. Product weights were chosen after earlier results were observed; this sensitivity view is not a trained-model improvement.
|
| 2 |
|
| 3 |
| Model | Current 30/25/15/15/15 | Earlier 25/25/15/15/20 | Original four-panel mean |
|
| 4 |
|---|---:|---:|---:|
|
|
|
|
| 5 |
| Lux | 76.72 | 76.43 | 79.06 |
|
| 6 |
+
| Nox | 73.09 | 72.42 | 75.03 |
|
| 7 |
| Kev-9B | 71.89 | 72.01 | 73.19 |
|
| 8 |
| Kev-4B | 70.09 | 70.30 | 71.73 |
|
| 9 |
| Qwen3.5-9B | 69.73 | 69.70 | 71.99 |
|
|
|
|
| 14 |
| Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
|
| 15 |
| Laya · English | 51.03 | 50.85 | 51.76 |
|
| 16 |
| Laya · Multilingual | 47.19 | 47.18 | 48.56 |
|
| 17 |
+
| Jev | 81.05 | 81.45 | 82.45 |
|
TASKS.md
CHANGED
|
@@ -5,93 +5,93 @@ Accuracy (%) on the same requested rows. Bold marks a Decision-family result str
|
|
| 5 |
<details>
|
| 6 |
<summary>Decisions · 10 tasks</summary>
|
| 7 |
|
| 8 |
-
| Task | n |
|
| 9 |
-
|
|
| 10 |
-
| News classification | 128 |
|
| 11 |
-
| Boolean constraints | 64 |
|
| 12 |
-
| Entity classification | 112 | 96.43 | 96.43 |
|
| 13 |
-
| Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 |
|
| 14 |
-
| Evidence placement | 96 |
|
| 15 |
-
| Ordered rubric | 64 | 100.00 |
|
| 16 |
-
| Relation composition | 96 |
|
| 17 |
-
| Scoped evidence | 96 |
|
| 18 |
-
| State tracking | 96 | 33.33 |
|
| 19 |
-
| In / out of menu | 64 |
|
| 20 |
|
| 21 |
</details>
|
| 22 |
|
| 23 |
<details>
|
| 24 |
<summary>Composition · 10 tasks</summary>
|
| 25 |
|
| 26 |
-
| Task | n |
|
| 27 |
-
|
|
| 28 |
-
| Record identity | 80 |
|
| 29 |
-
| Capacity assignment | 80 |
|
| 30 |
-
| Constraint assignment | 80 |
|
| 31 |
-
| Intent routing · EN | 120 |
|
| 32 |
-
| Intent routing · ZH | 120 |
|
| 33 |
-
| Multiset reconciliation | 80 |
|
| 34 |
-
| Ordered service loss | 80 |
|
| 35 |
-
| Conflicting rule closure | 80 |
|
| 36 |
-
| Temporal exclusion | 80 |
|
| 37 |
-
| Transaction recovery | 80 |
|
| 38 |
|
| 39 |
</details>
|
| 40 |
|
| 41 |
<details>
|
| 42 |
<summary>Reading · 3 tasks</summary>
|
| 43 |
|
| 44 |
-
| Task | n |
|
| 45 |
-
|
|
| 46 |
-
| Yes / no reading | 160 |
|
| 47 |
-
| Reading · EN | 160 |
|
| 48 |
-
| Reading · ZH | 160 |
|
| 49 |
|
| 50 |
</details>
|
| 51 |
|
| 52 |
<details>
|
| 53 |
<summary>Inference · 4 tasks</summary>
|
| 54 |
|
| 55 |
-
| Task | n |
|
| 56 |
-
|
|
| 57 |
-
| Contextual reasoning | 120 |
|
| 58 |
-
| Answerability | 120 |
|
| 59 |
-
| Textual entailment | 120 |
|
| 60 |
-
| Scientific inference | 120 |
|
| 61 |
|
| 62 |
</details>
|
| 63 |
|
| 64 |
<details>
|
| 65 |
<summary>Transfer · 27 tasks</summary>
|
| 66 |
|
| 67 |
-
| Task | n |
|
| 68 |
-
|
|
| 69 |
-
| Buried emotion | 20 |
|
| 70 |
-
| Buried paraphrase | 20 |
|
| 71 |
-
| Buried entailment | 20 |
|
| 72 |
-
| Buried offensive-language detection | 20 |
|
| 73 |
-
| Combined policy conditions | 32 |
|
| 74 |
-
| Policy exceptions | 32 |
|
| 75 |
-
| Policy negation | 32 |
|
| 76 |
-
| Authorization contrast | 40 | 100.00 |
|
| 77 |
-
| Deadline contrast | 40 |
|
| 78 |
-
| Emotion | 80 |
|
| 79 |
-
| MMLU | 80 |
|
| 80 |
-
| MMLU-Pro | 200 |
|
| 81 |
-
| Paraphrase | 80 |
|
| 82 |
-
| Question entailment | 80 |
|
| 83 |
-
| Science questions | 80 | 100.00 |
|
| 84 |
-
| Offensive-language detection | 80 |
|
| 85 |
-
| Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 |
|
| 86 |
-
| Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 |
|
| 87 |
-
| Evidence control · deadline | 10 |
|
| 88 |
-
| Evidence control · late fee | 10 |
|
| 89 |
-
| Evidence control · quantity limit | 10 | 100.00 |
|
| 90 |
-
| Evidence control · return window | 10 |
|
| 91 |
-
| Evidence control · shipping delay | 10 | 90.00 |
|
| 92 |
-
| Evidence control · sla response | 10 |
|
| 93 |
-
| Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 |
|
| 94 |
-
| Evidence control · volume discount | 10 | 100.00 |
|
| 95 |
-
| Evidence control · warranty claim | 10 |
|
| 96 |
|
| 97 |
</details>
|
|
|
|
| 5 |
<details>
|
| 6 |
<summary>Decisions · 10 tasks</summary>
|
| 7 |
|
| 8 |
+
| Task | n | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Jev |
|
| 9 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 10 |
+
| News classification | 128 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 | 85.16 |
|
| 11 |
+
| Boolean constraints | 64 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 | 100.00 |
|
| 12 |
+
| Entity classification | 112 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 | 96.43 |
|
| 13 |
+
| Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 | 100.00 |
|
| 14 |
+
| Evidence placement | 96 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 | 30.21 |
|
| 15 |
+
| Ordered rubric | 64 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 | 100.00 |
|
| 16 |
+
| Relation composition | 96 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 | 56.25 |
|
| 17 |
+
| Scoped evidence | 96 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 | 89.58 |
|
| 18 |
+
| State tracking | 96 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 | 33.33 |
|
| 19 |
+
| In / out of menu | 64 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 | 100.00 |
|
| 20 |
|
| 21 |
</details>
|
| 22 |
|
| 23 |
<details>
|
| 24 |
<summary>Composition · 10 tasks</summary>
|
| 25 |
|
| 26 |
+
| Task | n | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Jev |
|
| 27 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 28 |
+
| Record identity | 80 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 | 76.25 |
|
| 29 |
+
| Capacity assignment | 80 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 | 68.75 |
|
| 30 |
+
| Constraint assignment | 80 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 | 55.00 |
|
| 31 |
+
| Intent routing · EN | 120 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 | 91.67 |
|
| 32 |
+
| Intent routing · ZH | 120 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 | 88.33 |
|
| 33 |
+
| Multiset reconciliation | 80 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 | 58.75 |
|
| 34 |
+
| Ordered service loss | 80 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 | 42.50 |
|
| 35 |
+
| Conflicting rule closure | 80 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 | 66.25 |
|
| 36 |
+
| Temporal exclusion | 80 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 | 41.25 |
|
| 37 |
+
| Transaction recovery | 80 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 | 75.00 |
|
| 38 |
|
| 39 |
</details>
|
| 40 |
|
| 41 |
<details>
|
| 42 |
<summary>Reading · 3 tasks</summary>
|
| 43 |
|
| 44 |
+
| Task | n | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Jev |
|
| 45 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 46 |
+
| Yes / no reading | 160 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 | 92.50 |
|
| 47 |
+
| Reading · EN | 160 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 | 96.88 |
|
| 48 |
+
| Reading · ZH | 160 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 | 96.25 |
|
| 49 |
|
| 50 |
</details>
|
| 51 |
|
| 52 |
<details>
|
| 53 |
<summary>Inference · 4 tasks</summary>
|
| 54 |
|
| 55 |
+
| Task | n | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Jev |
|
| 56 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 57 |
+
| Contextual reasoning | 120 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 | 86.67 |
|
| 58 |
+
| Answerability | 120 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 | 90.83 |
|
| 59 |
+
| Textual entailment | 120 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 | 82.50 |
|
| 60 |
+
| Scientific inference | 120 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 | 99.17 |
|
| 61 |
|
| 62 |
</details>
|
| 63 |
|
| 64 |
<details>
|
| 65 |
<summary>Transfer · 27 tasks</summary>
|
| 66 |
|
| 67 |
+
| Task | n | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Jev |
|
| 68 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 69 |
+
| Buried emotion | 20 | 40.00 | 25.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 | 70.00 |
|
| 70 |
+
| Buried paraphrase | 20 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 | 95.00 |
|
| 71 |
+
| Buried entailment | 20 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 | 95.00 |
|
| 72 |
+
| Buried offensive-language detection | 20 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 | 55.00 |
|
| 73 |
+
| Combined policy conditions | 32 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 | 96.88 |
|
| 74 |
+
| Policy exceptions | 32 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 | 100.00 |
|
| 75 |
+
| Policy negation | 32 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 | 90.62 |
|
| 76 |
+
| Authorization contrast | 40 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 | 100.00 |
|
| 77 |
+
| Deadline contrast | 40 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 | 92.50 |
|
| 78 |
+
| Emotion | 80 | 61.25 | 60.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 | 67.50 |
|
| 79 |
+
| MMLU | 80 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 | 88.75 |
|
| 80 |
+
| MMLU-Pro | 200 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 | 84.00 |
|
| 81 |
+
| Paraphrase | 80 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 | 87.50 |
|
| 82 |
+
| Question entailment | 80 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 | 91.25 |
|
| 83 |
+
| Science questions | 80 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 | 100.00 |
|
| 84 |
+
| Offensive-language detection | 80 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 | 76.25 |
|
| 85 |
+
| Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 | 100.00 |
|
| 86 |
+
| Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 | 100.00 |
|
| 87 |
+
| Evidence control · deadline | 10 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 | 100.00 |
|
| 88 |
+
| Evidence control · late fee | 10 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 | 100.00 |
|
| 89 |
+
| Evidence control · quantity limit | 10 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 | 100.00 |
|
| 90 |
+
| Evidence control · return window | 10 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 | 50.00 |
|
| 91 |
+
| Evidence control · shipping delay | 10 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 | 90.00 |
|
| 92 |
+
| Evidence control · sla response | 10 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 | 100.00 |
|
| 93 |
+
| Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 | 100.00 |
|
| 94 |
+
| Evidence control · volume discount | 10 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 | 100.00 |
|
| 95 |
+
| Evidence control · warranty claim | 10 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 | 90.00 |
|
| 96 |
|
| 97 |
</details>
|
USAGE.md
CHANGED
|
@@ -49,12 +49,14 @@ The model card must publish output from that model's real run; this document inv
|
|
| 49 |
|
| 50 |
| Type | Input criteria | Answer |
|
| 51 |
|---|---|---|
|
| 52 |
-
| Choice | Ordered mapping of 2–255 external IDs to
|
| 53 |
| Noul | Optional mapping containing only `false` / `true` descriptions | `noul`: P(true); a hard judgment uses `>= 0.5` |
|
| 54 |
| Score | Ordered list of 2–10 rubric descriptions | `probabilities` over string indices, expected level index `score`, `legend`, `confidence` |
|
| 55 |
|
| 56 |
Response shape is `{"model": name, "answers": {question_name: answer}, "usage": {"input_tokens": total, "scored_questions": count}}`. Choice ties select the earliest candidate in insertion order. Score returns the expected ordinal **index**, not an arbitrary supplied numeric value; this adapter does not implement a supplied-values extension. Noul 0.5 is interpreted as true. Confidence is `(K * max(p) - 1)/(K - 1)`, clipped to [0,1], and is not a claimed reproduction of Jev's confidence statistic. Shipped temperature calibration does not make every confidence value a correctness guarantee.
|
| 57 |
|
|
|
|
|
|
|
| 58 |
Question names are preserved as opaque bookkeeping IDs; candidate IDs and descriptions use the frozen renderer. Native objects are deterministically serialized with sorted JSON keys; strings preserve their contents. Do not interpret object/string field-order differences as identical token inputs. Non-finite or non-JSON inputs are rejected.
|
| 59 |
|
| 60 |
The bound is **16,384 tokens per complete question**, including its state, instructions, all candidates and readout suffix. Questions are separate sequences, grouped in fixed batches of eight; the state is repeated for each question and counted repeatedly in `usage.input_tokens`. All questions are encoded before any forward pass. If any exceeds the bound, the whole call raises `ValueError` without truncation or partial answers. A lower `max_length` can be chosen at load time; a higher limit is rejected. This native wrapper does not impose the Studio's separate 16-question UI limit.
|
|
|
|
| 49 |
|
| 50 |
| Type | Input criteria | Answer |
|
| 51 |
|---|---|---|
|
| 52 |
+
| Choice | Ordered mapping of 2–255 external IDs to descriptions or null | `probabilities`, selected `choice` ID, `confidence` |
|
| 53 |
| Noul | Optional mapping containing only `false` / `true` descriptions | `noul`: P(true); a hard judgment uses `>= 0.5` |
|
| 54 |
| Score | Ordered list of 2–10 rubric descriptions | `probabilities` over string indices, expected level index `score`, `legend`, `confidence` |
|
| 55 |
|
| 56 |
Response shape is `{"model": name, "answers": {question_name: answer}, "usage": {"input_tokens": total, "scored_questions": count}}`. Choice ties select the earliest candidate in insertion order. Score returns the expected ordinal **index**, not an arbitrary supplied numeric value; this adapter does not implement a supplied-values extension. Noul 0.5 is interpreted as true. Confidence is `(K * max(p) - 1)/(K - 1)`, clipped to [0,1], and is not a claimed reproduction of Jev's confidence statistic. Shipped temperature calibration does not make every confidence value a correctness guarantee.
|
| 57 |
|
| 58 |
+
For Choice, a null description uses the original candidate ID as its semantic description. Explicit descriptions, including an empty string, are preserved. The original request object is not mutated. Noul and Score are unchanged.
|
| 59 |
+
|
| 60 |
Question names are preserved as opaque bookkeeping IDs; candidate IDs and descriptions use the frozen renderer. Native objects are deterministically serialized with sorted JSON keys; strings preserve their contents. Do not interpret object/string field-order differences as identical token inputs. Non-finite or non-JSON inputs are rejected.
|
| 61 |
|
| 62 |
The bound is **16,384 tokens per complete question**, including its state, instructions, all candidates and readout suffix. Questions are separate sequences, grouped in fixed batches of eight; the state is repeated for each question and counted repeatedly in `usage.input_tokens`. All questions are encoded before any forward pass. If any exceeds the bound, the whole call raises `ValueError` without truncation or partial answers. A lower `max_length` can be chosen at load time; a higher limit is rejected. This native wrapper does not impose the Studio's separate 16-question UI limit.
|
WEIGHTING.md
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
The same latest model predictions are shown under three weighting schemes. Product weights were chosen after earlier results were observed; this sensitivity view is not a trained-model improvement.
|
| 2 |
+
|
| 3 |
+
| Model | Current 30/25/15/15/15 | Earlier 25/25/15/15/20 | Original four-panel mean |
|
| 4 |
+
|---|---:|---:|---:|
|
| 5 |
+
| Lux | 76.72 | 76.43 | 79.06 |
|
| 6 |
+
| Nox | 73.09 | 72.42 | 75.03 |
|
| 7 |
+
| Kev-9B | 71.89 | 72.01 | 73.19 |
|
| 8 |
+
| Kev-4B | 70.09 | 70.30 | 71.73 |
|
| 9 |
+
| Qwen3.5-9B | 69.73 | 69.70 | 71.99 |
|
| 10 |
+
| Decider | 67.71 | 67.97 | 71.75 |
|
| 11 |
+
| Qwen3.5-4B | 67.29 | 67.24 | 70.25 |
|
| 12 |
+
| Sol | 66.32 | 65.48 | 70.14 |
|
| 13 |
+
| Kev-0.8B | 58.28 | 58.33 | 59.75 |
|
| 14 |
+
| Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
|
| 15 |
+
| Laya · English | 51.03 | 50.85 | 51.76 |
|
| 16 |
+
| Laya · Multilingual | 47.19 | 47.18 | 48.56 |
|
| 17 |
+
| Jev | 81.05 | 81.45 | 82.45 |
|
assets/decision-matrix.pdf
CHANGED
|
Binary files a/assets/decision-matrix.pdf and b/assets/decision-matrix.pdf differ
|
|
|
assets/decision-matrix.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
assets/decision-matrix.svg
CHANGED
|
|
|
|
assets/decision-ranking.pdf
CHANGED
|
Binary files a/assets/decision-ranking.pdf and b/assets/decision-ranking.pdf differ
|
|
|
assets/decision-ranking.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
assets/decision-ranking.svg
CHANGED
|
|
|
|
bundle-manifest.json
CHANGED
|
@@ -1,12 +1,17 @@
|
|
| 1 |
{
|
| 2 |
"format": "research-pointer-bundle-v1",
|
| 3 |
-
"status": "
|
| 4 |
"files": [
|
| 5 |
{
|
| 6 |
"file": "NORMALIZATION_RUNTIME.md",
|
| 7 |
"bytes": 756,
|
| 8 |
"sha256": "cf8e6ce1f07687a68b6adeb98e6704b8f292cb6adc7b10e4b84e8e70aa5472e5"
|
| 9 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
{
|
| 11 |
"file": "RUNTIME_BINDING.json",
|
| 12 |
"bytes": 2621,
|
|
@@ -54,8 +59,8 @@
|
|
| 54 |
},
|
| 55 |
{
|
| 56 |
"file": "code/decision_api.py",
|
| 57 |
-
"bytes":
|
| 58 |
-
"sha256": "
|
| 59 |
},
|
| 60 |
{
|
| 61 |
"file": "code/decision_model.py",
|
|
@@ -221,7 +226,7 @@
|
|
| 221 |
}
|
| 222 |
],
|
| 223 |
"source_model_code_sha256": "d3e28489c09f3bd7130e2d43d92e0b5c4a08e25b09b21303904defb0ff1c3646",
|
| 224 |
-
"source_api_code_sha256": "
|
| 225 |
"dev_sha256": "45b4cd46acc4be53b95c7ed8f333b3533972296a4d0557c7b8c48c9cab6ced61",
|
| 226 |
"production_predictions_sha256": "ba311f5a6bc73182651bdf5c9744f0ed95921110501ddc93dcfd20a1d06bceff",
|
| 227 |
"temperature_sha256": "69d80e5b215e2c1e6872f146fdb7ea5aa95c9fd678ef96208960d50e4e7b635d",
|
|
@@ -230,7 +235,7 @@
|
|
| 230 |
"input_length_limit": 16384,
|
| 231 |
"original_checkpoint_name": "winner",
|
| 232 |
"no_publication_performed": true,
|
| 233 |
-
"source_bundle_manifest_sha256": "
|
| 234 |
"normalization_profile_sha256": "be32858d15233e0a3fbee0e4257fb02be0b3439deee4eb9c3f61151df7b73850",
|
| 235 |
"public_wrapper_included": true,
|
| 236 |
"source_published_bundle_manifest_sha256": "92d7f5be5e1ef21ee682f574b35cc01de4dd8ab16b8aa6edf0dfbdfcfcfabba3",
|
|
|
|
| 1 |
{
|
| 2 |
"format": "research-pointer-bundle-v1",
|
| 3 |
+
"status": "null-description-candidate-awaiting-public-proof-and-full-regression",
|
| 4 |
"files": [
|
| 5 |
{
|
| 6 |
"file": "NORMALIZATION_RUNTIME.md",
|
| 7 |
"bytes": 756,
|
| 8 |
"sha256": "cf8e6ce1f07687a68b6adeb98e6704b8f292cb6adc7b10e4b84e8e70aa5472e5"
|
| 9 |
},
|
| 10 |
+
{
|
| 11 |
+
"file": "NULL_DESCRIPTION_RENDERING.json",
|
| 12 |
+
"bytes": 514,
|
| 13 |
+
"sha256": "3b531cab60cba35648ab71fe4be7fd1af619f02ee337e3beed6296fdfc8b9ec8"
|
| 14 |
+
},
|
| 15 |
{
|
| 16 |
"file": "RUNTIME_BINDING.json",
|
| 17 |
"bytes": 2621,
|
|
|
|
| 59 |
},
|
| 60 |
{
|
| 61 |
"file": "code/decision_api.py",
|
| 62 |
+
"bytes": 10978,
|
| 63 |
+
"sha256": "6716bdabca3cd2f1aa447d42a6cf53ee62d7ea97984a01f16b4c09fac8aaf17e"
|
| 64 |
},
|
| 65 |
{
|
| 66 |
"file": "code/decision_model.py",
|
|
|
|
| 226 |
}
|
| 227 |
],
|
| 228 |
"source_model_code_sha256": "d3e28489c09f3bd7130e2d43d92e0b5c4a08e25b09b21303904defb0ff1c3646",
|
| 229 |
+
"source_api_code_sha256": "6716bdabca3cd2f1aa447d42a6cf53ee62d7ea97984a01f16b4c09fac8aaf17e",
|
| 230 |
"dev_sha256": "45b4cd46acc4be53b95c7ed8f333b3533972296a4d0557c7b8c48c9cab6ced61",
|
| 231 |
"production_predictions_sha256": "ba311f5a6bc73182651bdf5c9744f0ed95921110501ddc93dcfd20a1d06bceff",
|
| 232 |
"temperature_sha256": "69d80e5b215e2c1e6872f146fdb7ea5aa95c9fd678ef96208960d50e4e7b635d",
|
|
|
|
| 235 |
"input_length_limit": 16384,
|
| 236 |
"original_checkpoint_name": "winner",
|
| 237 |
"no_publication_performed": true,
|
| 238 |
+
"source_bundle_manifest_sha256": "83876db506b2d98e3e8ce7d34310b21f97bac30d4f5aef3371053798bff08830",
|
| 239 |
"normalization_profile_sha256": "be32858d15233e0a3fbee0e4257fb02be0b3439deee4eb9c3f61151df7b73850",
|
| 240 |
"public_wrapper_included": true,
|
| 241 |
"source_published_bundle_manifest_sha256": "92d7f5be5e1ef21ee682f574b35cc01de4dd8ab16b8aa6edf0dfbdfcfcfabba3",
|
code/decision_api.py
CHANGED
|
@@ -55,7 +55,7 @@ def question_row(state, name, question):
|
|
| 55 |
if not isinstance(criteria,dict) or not 2<=len(criteria)<=255:
|
| 56 |
raise ValueError('choice requires a mapping of 2..255 criteria')
|
| 57 |
if not all(isinstance(k,str) for k in criteria):raise ValueError('Choice keys must be strings')
|
| 58 |
-
options=[{'key':key,'description':value} for key,value in criteria.items()]
|
| 59 |
# The question name is used for bookkeeping only; encoders never render id.
|
| 60 |
return {'id':name,'state':state,'instructions':question['instructions'],
|
| 61 |
'options':options,'task_type':kind,'family':'inference'}
|
|
|
|
| 55 |
if not isinstance(criteria,dict) or not 2<=len(criteria)<=255:
|
| 56 |
raise ValueError('choice requires a mapping of 2..255 criteria')
|
| 57 |
if not all(isinstance(k,str) for k in criteria):raise ValueError('Choice keys must be strings')
|
| 58 |
+
options=[{'key':key,'description':key if value is None else value} for key,value in criteria.items()]
|
| 59 |
# The question name is used for bookkeeping only; encoders never render id.
|
| 60 |
return {'id':name,'state':state,'instructions':question['instructions'],
|
| 61 |
'options':options,'task_type':kind,'family':'inference'}
|
metrics/benchmark.json
CHANGED
|
@@ -9,7 +9,7 @@
|
|
| 9 |
},
|
| 10 |
"metric_design": "User-requested, outcome-informed decision-priority weights; observed regression data, not blind testing.",
|
| 11 |
"protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
|
| 12 |
-
"statistics_sha256": "
|
| 13 |
"models": {
|
| 14 |
"Jev": {
|
| 15 |
"panels": {
|
|
@@ -10330,11 +10330,11 @@
|
|
| 10330 |
]
|
| 10331 |
},
|
| 10332 |
"transfer_v9_test": {
|
| 10333 |
-
"point": 0.
|
| 10334 |
-
"exact_fraction": "
|
| 10335 |
"ci95": [
|
| 10336 |
-
0.
|
| 10337 |
-
0.
|
| 10338 |
]
|
| 10339 |
},
|
| 10340 |
"original_four_panel_mean": {
|
|
@@ -10346,11 +10346,11 @@
|
|
| 10346 |
]
|
| 10347 |
},
|
| 10348 |
"weighted_mean": {
|
| 10349 |
-
"point": 0.
|
| 10350 |
-
"exact_fraction": "
|
| 10351 |
"ci95": [
|
| 10352 |
-
0.
|
| 10353 |
-
0.
|
| 10354 |
]
|
| 10355 |
}
|
| 10356 |
},
|
|
@@ -10496,50 +10496,50 @@
|
|
| 10496 |
"clean": {
|
| 10497 |
"requested": 1046,
|
| 10498 |
"valid_probability_rows": 1046,
|
| 10499 |
-
"correct":
|
| 10500 |
-
"accuracy": 0.
|
| 10501 |
"full_probability_metrics": {
|
| 10502 |
"n": 1046,
|
| 10503 |
-
"nll": 0.
|
| 10504 |
-
"acc": 0.
|
| 10505 |
-
"ece": 0.
|
| 10506 |
-
"brier": 0.
|
| 10507 |
-
"mean_conf": 0.
|
| 10508 |
-
"confident_error_rate": 0.
|
| 10509 |
-
"coverage_at_0_9": 0.
|
| 10510 |
-
"accuracy_at_0_9": 0.
|
| 10511 |
-
"coverage_at_5pct_error": 0.
|
| 10512 |
-
"coverage_at_1pct_error": 0.
|
| 10513 |
-
"aurc": 0.
|
| 10514 |
-
"error_rate_at_0_9": 0.
|
| 10515 |
-
"confidence_bias": 0.
|
| 10516 |
"top_bins": {
|
| 10517 |
"0.9": {
|
| 10518 |
-
"n":
|
| 10519 |
-
"errors":
|
| 10520 |
-
"error_rate": 0.
|
| 10521 |
},
|
| 10522 |
"0.95": {
|
| 10523 |
-
"n":
|
| 10524 |
-
"errors":
|
| 10525 |
-
"error_rate": 0.
|
| 10526 |
},
|
| 10527 |
"0.99": {
|
| 10528 |
-
"n":
|
| 10529 |
-
"errors":
|
| 10530 |
-
"error_rate": 0.
|
| 10531 |
}
|
| 10532 |
},
|
| 10533 |
"selective": {
|
| 10534 |
"0.5": {
|
| 10535 |
"coverage": 0.5,
|
| 10536 |
-
"accuracy": 0.
|
| 10537 |
-
"confidence_cutoff": 0.
|
| 10538 |
},
|
| 10539 |
"0.8": {
|
| 10540 |
"coverage": 0.8001912045889101,
|
| 10541 |
-
"accuracy": 0.
|
| 10542 |
-
"confidence_cutoff": 0.
|
| 10543 |
}
|
| 10544 |
},
|
| 10545 |
"score_mae": 0.35059676214309027,
|
|
@@ -10547,46 +10547,46 @@
|
|
| 10547 |
},
|
| 10548 |
"covered_only_probability_metrics": {
|
| 10549 |
"n": 1046,
|
| 10550 |
-
"nll": 0.
|
| 10551 |
-
"acc": 0.
|
| 10552 |
-
"ece": 0.
|
| 10553 |
-
"brier": 0.
|
| 10554 |
-
"mean_conf": 0.
|
| 10555 |
-
"confident_error_rate": 0.
|
| 10556 |
-
"coverage_at_0_9": 0.
|
| 10557 |
-
"accuracy_at_0_9": 0.
|
| 10558 |
-
"coverage_at_5pct_error": 0.
|
| 10559 |
-
"coverage_at_1pct_error": 0.
|
| 10560 |
-
"aurc": 0.
|
| 10561 |
-
"error_rate_at_0_9": 0.
|
| 10562 |
-
"confidence_bias": 0.
|
| 10563 |
"top_bins": {
|
| 10564 |
"0.9": {
|
| 10565 |
-
"n":
|
| 10566 |
-
"errors":
|
| 10567 |
-
"error_rate": 0.
|
| 10568 |
},
|
| 10569 |
"0.95": {
|
| 10570 |
-
"n":
|
| 10571 |
-
"errors":
|
| 10572 |
-
"error_rate": 0.
|
| 10573 |
},
|
| 10574 |
"0.99": {
|
| 10575 |
-
"n":
|
| 10576 |
-
"errors":
|
| 10577 |
-
"error_rate": 0.
|
| 10578 |
}
|
| 10579 |
},
|
| 10580 |
"selective": {
|
| 10581 |
"0.5": {
|
| 10582 |
"coverage": 0.5,
|
| 10583 |
-
"accuracy": 0.
|
| 10584 |
-
"confidence_cutoff": 0.
|
| 10585 |
},
|
| 10586 |
"0.8": {
|
| 10587 |
"coverage": 0.8001912045889101,
|
| 10588 |
-
"accuracy": 0.
|
| 10589 |
-
"confidence_cutoff": 0.
|
| 10590 |
}
|
| 10591 |
},
|
| 10592 |
"score_mae": 0.35059676214309027,
|
|
@@ -10598,33 +10598,33 @@
|
|
| 10598 |
"buried_emotion": {
|
| 10599 |
"requested": 20,
|
| 10600 |
"valid_probability_rows": 20,
|
| 10601 |
-
"correct":
|
| 10602 |
-
"accuracy": 0.
|
| 10603 |
"full_probability_metrics": {
|
| 10604 |
"n": 20,
|
| 10605 |
-
"nll": 1.
|
| 10606 |
-
"acc": 0.
|
| 10607 |
-
"ece": 0.
|
| 10608 |
-
"brier": 0.
|
| 10609 |
-
"mean_conf": 0.
|
| 10610 |
"confident_error_rate": 0.0,
|
| 10611 |
-
"coverage_at_0_9": 0.
|
| 10612 |
-
"accuracy_at_0_9":
|
| 10613 |
-
"coverage_at_5pct_error": 0.
|
| 10614 |
-
"coverage_at_1pct_error": 0.
|
| 10615 |
-
"aurc": 0.
|
| 10616 |
-
"error_rate_at_0_9":
|
| 10617 |
-
"confidence_bias": 0.
|
| 10618 |
"top_bins": {
|
| 10619 |
"0.9": {
|
| 10620 |
-
"n":
|
| 10621 |
"errors": 0,
|
| 10622 |
-
"error_rate":
|
| 10623 |
},
|
| 10624 |
"0.95": {
|
| 10625 |
-
"n":
|
| 10626 |
"errors": 0,
|
| 10627 |
-
"error_rate":
|
| 10628 |
},
|
| 10629 |
"0.99": {
|
| 10630 |
"n": 0,
|
|
@@ -10635,41 +10635,41 @@
|
|
| 10635 |
"selective": {
|
| 10636 |
"0.5": {
|
| 10637 |
"coverage": 0.5,
|
| 10638 |
-
"accuracy": 0.
|
| 10639 |
-
"confidence_cutoff": 0.
|
| 10640 |
},
|
| 10641 |
"0.8": {
|
| 10642 |
"coverage": 0.8,
|
| 10643 |
-
"accuracy": 0.
|
| 10644 |
-
"confidence_cutoff": 0.
|
| 10645 |
}
|
| 10646 |
}
|
| 10647 |
},
|
| 10648 |
"covered_only_probability_metrics": {
|
| 10649 |
"n": 20,
|
| 10650 |
-
"nll": 1.
|
| 10651 |
-
"acc": 0.
|
| 10652 |
-
"ece": 0.
|
| 10653 |
-
"brier": 0.
|
| 10654 |
-
"mean_conf": 0.
|
| 10655 |
"confident_error_rate": 0.0,
|
| 10656 |
-
"coverage_at_0_9": 0.
|
| 10657 |
-
"accuracy_at_0_9":
|
| 10658 |
-
"coverage_at_5pct_error": 0.
|
| 10659 |
-
"coverage_at_1pct_error": 0.
|
| 10660 |
-
"aurc": 0.
|
| 10661 |
-
"error_rate_at_0_9":
|
| 10662 |
-
"confidence_bias": 0.
|
| 10663 |
"top_bins": {
|
| 10664 |
"0.9": {
|
| 10665 |
-
"n":
|
| 10666 |
"errors": 0,
|
| 10667 |
-
"error_rate":
|
| 10668 |
},
|
| 10669 |
"0.95": {
|
| 10670 |
-
"n":
|
| 10671 |
"errors": 0,
|
| 10672 |
-
"error_rate":
|
| 10673 |
},
|
| 10674 |
"0.99": {
|
| 10675 |
"n": 0,
|
|
@@ -10680,13 +10680,13 @@
|
|
| 10680 |
"selective": {
|
| 10681 |
"0.5": {
|
| 10682 |
"coverage": 0.5,
|
| 10683 |
-
"accuracy": 0.
|
| 10684 |
-
"confidence_cutoff": 0.
|
| 10685 |
},
|
| 10686 |
"0.8": {
|
| 10687 |
"coverage": 0.8,
|
| 10688 |
-
"accuracy": 0.
|
| 10689 |
-
"confidence_cutoff": 0.
|
| 10690 |
}
|
| 10691 |
}
|
| 10692 |
},
|
|
@@ -11475,95 +11475,95 @@
|
|
| 11475 |
"emotion": {
|
| 11476 |
"requested": 80,
|
| 11477 |
"valid_probability_rows": 80,
|
| 11478 |
-
"correct":
|
| 11479 |
-
"accuracy": 0.
|
| 11480 |
"full_probability_metrics": {
|
| 11481 |
"n": 80,
|
| 11482 |
-
"nll": 1.
|
| 11483 |
-
"acc": 0.
|
| 11484 |
-
"ece": 0.
|
| 11485 |
-
"brier": 0.
|
| 11486 |
-
"mean_conf": 0.
|
| 11487 |
-
"confident_error_rate": 0.
|
| 11488 |
-
"coverage_at_0_9": 0.
|
| 11489 |
-
"accuracy_at_0_9":
|
| 11490 |
-
"coverage_at_5pct_error": 0.
|
| 11491 |
-
"coverage_at_1pct_error": 0.
|
| 11492 |
-
"aurc": 0.
|
| 11493 |
-
"error_rate_at_0_9":
|
| 11494 |
-
"confidence_bias":
|
| 11495 |
"top_bins": {
|
| 11496 |
"0.9": {
|
| 11497 |
-
"n":
|
| 11498 |
-
"errors":
|
| 11499 |
-
"error_rate":
|
| 11500 |
},
|
| 11501 |
"0.95": {
|
| 11502 |
-
"n":
|
| 11503 |
-
"errors":
|
| 11504 |
-
"error_rate":
|
| 11505 |
},
|
| 11506 |
"0.99": {
|
| 11507 |
-
"n":
|
| 11508 |
-
"errors":
|
| 11509 |
-
"error_rate":
|
| 11510 |
}
|
| 11511 |
},
|
| 11512 |
"selective": {
|
| 11513 |
"0.5": {
|
| 11514 |
"coverage": 0.5,
|
| 11515 |
-
"accuracy": 0.
|
| 11516 |
-
"confidence_cutoff": 0.
|
| 11517 |
},
|
| 11518 |
"0.8": {
|
| 11519 |
"coverage": 0.8,
|
| 11520 |
-
"accuracy": 0.
|
| 11521 |
-
"confidence_cutoff": 0.
|
| 11522 |
}
|
| 11523 |
}
|
| 11524 |
},
|
| 11525 |
"covered_only_probability_metrics": {
|
| 11526 |
"n": 80,
|
| 11527 |
-
"nll": 1.
|
| 11528 |
-
"acc": 0.
|
| 11529 |
-
"ece": 0.
|
| 11530 |
-
"brier": 0.
|
| 11531 |
-
"mean_conf": 0.
|
| 11532 |
-
"confident_error_rate": 0.
|
| 11533 |
-
"coverage_at_0_9": 0.
|
| 11534 |
-
"accuracy_at_0_9":
|
| 11535 |
-
"coverage_at_5pct_error": 0.
|
| 11536 |
-
"coverage_at_1pct_error": 0.
|
| 11537 |
-
"aurc": 0.
|
| 11538 |
-
"error_rate_at_0_9":
|
| 11539 |
-
"confidence_bias":
|
| 11540 |
"top_bins": {
|
| 11541 |
"0.9": {
|
| 11542 |
-
"n":
|
| 11543 |
-
"errors":
|
| 11544 |
-
"error_rate":
|
| 11545 |
},
|
| 11546 |
"0.95": {
|
| 11547 |
-
"n":
|
| 11548 |
-
"errors":
|
| 11549 |
-
"error_rate":
|
| 11550 |
},
|
| 11551 |
"0.99": {
|
| 11552 |
-
"n":
|
| 11553 |
-
"errors":
|
| 11554 |
-
"error_rate":
|
| 11555 |
}
|
| 11556 |
},
|
| 11557 |
"selective": {
|
| 11558 |
"0.5": {
|
| 11559 |
"coverage": 0.5,
|
| 11560 |
-
"accuracy": 0.
|
| 11561 |
-
"confidence_cutoff": 0.
|
| 11562 |
},
|
| 11563 |
"0.8": {
|
| 11564 |
"coverage": 0.8,
|
| 11565 |
-
"accuracy": 0.
|
| 11566 |
-
"confidence_cutoff": 0.
|
| 11567 |
}
|
| 11568 |
}
|
| 11569 |
},
|
|
@@ -13243,14 +13243,14 @@
|
|
| 13243 |
"coverage": 1.0
|
| 13244 |
}
|
| 13245 |
},
|
| 13246 |
-
"clean_task_macro_accuracy": 0.
|
| 13247 |
"permutation": {
|
| 13248 |
"requested_pairs": 36,
|
| 13249 |
"valid_pairs": 36,
|
| 13250 |
-
"both_correct_requested": 0.
|
| 13251 |
-
"flip_rate": 0.
|
| 13252 |
-
"covered_only_flip_rate": 0.
|
| 13253 |
-
"half_l1": 0.
|
| 13254 |
},
|
| 13255 |
"unknowable": {
|
| 13256 |
"requested": 110,
|
|
@@ -13378,7 +13378,7 @@
|
|
| 13378 |
"truncated_questions": 0
|
| 13379 |
},
|
| 13380 |
"complete_unchanged_upstream_report": {
|
| 13381 |
-
"objective": -0.
|
| 13382 |
"paired_flip": {
|
| 13383 |
"pairs": 64,
|
| 13384 |
"flip_rate": 0.828125,
|
|
@@ -13399,46 +13399,46 @@
|
|
| 13399 |
},
|
| 13400 |
"clean": {
|
| 13401 |
"n": 1046,
|
| 13402 |
-
"nll": 0.
|
| 13403 |
-
"acc": 0.
|
| 13404 |
-
"ece": 0.
|
| 13405 |
-
"brier": 0.
|
| 13406 |
-
"mean_conf": 0.
|
| 13407 |
-
"confident_error_rate": 0.
|
| 13408 |
-
"coverage_at_0_9": 0.
|
| 13409 |
-
"accuracy_at_0_9": 0.
|
| 13410 |
-
"coverage_at_5pct_error": 0.
|
| 13411 |
-
"coverage_at_1pct_error": 0.
|
| 13412 |
-
"aurc": 0.
|
| 13413 |
-
"error_rate_at_0_9": 0.
|
| 13414 |
-
"confidence_bias": 0.
|
| 13415 |
"top_bins": {
|
| 13416 |
"0.9": {
|
| 13417 |
-
"n":
|
| 13418 |
-
"errors":
|
| 13419 |
-
"error_rate": 0.
|
| 13420 |
},
|
| 13421 |
"0.95": {
|
| 13422 |
-
"n":
|
| 13423 |
-
"errors":
|
| 13424 |
-
"error_rate": 0.
|
| 13425 |
},
|
| 13426 |
"0.99": {
|
| 13427 |
-
"n":
|
| 13428 |
-
"errors":
|
| 13429 |
-
"error_rate": 0.
|
| 13430 |
}
|
| 13431 |
},
|
| 13432 |
"selective": {
|
| 13433 |
"0.5": {
|
| 13434 |
"coverage": 0.5,
|
| 13435 |
-
"accuracy": 0.
|
| 13436 |
-
"confidence_cutoff": 0.
|
| 13437 |
},
|
| 13438 |
"0.8": {
|
| 13439 |
"coverage": 0.8001912045889101,
|
| 13440 |
-
"accuracy": 0.
|
| 13441 |
-
"confidence_cutoff": 0.
|
| 13442 |
}
|
| 13443 |
},
|
| 13444 |
"score_mae": 0.35059676214309027,
|
|
@@ -13447,29 +13447,29 @@
|
|
| 13447 |
"tasks": {
|
| 13448 |
"buried_emotion": {
|
| 13449 |
"n": 20,
|
| 13450 |
-
"nll": 1.
|
| 13451 |
-
"acc": 0.
|
| 13452 |
-
"ece": 0.
|
| 13453 |
-
"brier": 0.
|
| 13454 |
-
"mean_conf": 0.
|
| 13455 |
"confident_error_rate": 0.0,
|
| 13456 |
-
"coverage_at_0_9": 0.
|
| 13457 |
-
"accuracy_at_0_9":
|
| 13458 |
-
"coverage_at_5pct_error": 0.
|
| 13459 |
-
"coverage_at_1pct_error": 0.
|
| 13460 |
-
"aurc": 0.
|
| 13461 |
-
"error_rate_at_0_9":
|
| 13462 |
-
"confidence_bias": 0.
|
| 13463 |
"top_bins": {
|
| 13464 |
"0.9": {
|
| 13465 |
-
"n":
|
| 13466 |
"errors": 0,
|
| 13467 |
-
"error_rate":
|
| 13468 |
},
|
| 13469 |
"0.95": {
|
| 13470 |
-
"n":
|
| 13471 |
"errors": 0,
|
| 13472 |
-
"error_rate":
|
| 13473 |
},
|
| 13474 |
"0.99": {
|
| 13475 |
"n": 0,
|
|
@@ -13480,13 +13480,13 @@
|
|
| 13480 |
"selective": {
|
| 13481 |
"0.5": {
|
| 13482 |
"coverage": 0.5,
|
| 13483 |
-
"accuracy": 0.
|
| 13484 |
-
"confidence_cutoff": 0.
|
| 13485 |
},
|
| 13486 |
"0.8": {
|
| 13487 |
"coverage": 0.8,
|
| 13488 |
-
"accuracy": 0.
|
| 13489 |
-
"confidence_cutoff": 0.
|
| 13490 |
}
|
| 13491 |
}
|
| 13492 |
},
|
|
@@ -13854,46 +13854,46 @@
|
|
| 13854 |
},
|
| 13855 |
"emotion": {
|
| 13856 |
"n": 80,
|
| 13857 |
-
"nll": 1.
|
| 13858 |
-
"acc": 0.
|
| 13859 |
-
"ece": 0.
|
| 13860 |
-
"brier": 0.
|
| 13861 |
-
"mean_conf": 0.
|
| 13862 |
-
"confident_error_rate": 0.
|
| 13863 |
-
"coverage_at_0_9": 0.
|
| 13864 |
-
"accuracy_at_0_9":
|
| 13865 |
-
"coverage_at_5pct_error": 0.
|
| 13866 |
-
"coverage_at_1pct_error": 0.
|
| 13867 |
-
"aurc": 0.
|
| 13868 |
-
"error_rate_at_0_9":
|
| 13869 |
-
"confidence_bias":
|
| 13870 |
"top_bins": {
|
| 13871 |
"0.9": {
|
| 13872 |
-
"n":
|
| 13873 |
-
"errors":
|
| 13874 |
-
"error_rate":
|
| 13875 |
},
|
| 13876 |
"0.95": {
|
| 13877 |
-
"n":
|
| 13878 |
-
"errors":
|
| 13879 |
-
"error_rate":
|
| 13880 |
},
|
| 13881 |
"0.99": {
|
| 13882 |
-
"n":
|
| 13883 |
-
"errors":
|
| 13884 |
-
"error_rate":
|
| 13885 |
}
|
| 13886 |
},
|
| 13887 |
"selective": {
|
| 13888 |
"0.5": {
|
| 13889 |
"coverage": 0.5,
|
| 13890 |
-
"accuracy": 0.
|
| 13891 |
-
"confidence_cutoff": 0.
|
| 13892 |
},
|
| 13893 |
"0.8": {
|
| 13894 |
"coverage": 0.8,
|
| 13895 |
-
"accuracy": 0.
|
| 13896 |
-
"confidence_cutoff": 0.
|
| 13897 |
}
|
| 13898 |
}
|
| 13899 |
},
|
|
@@ -15185,46 +15185,46 @@
|
|
| 15185 |
"variants": {
|
| 15186 |
"clean": {
|
| 15187 |
"n": 1156,
|
| 15188 |
-
"nll": 1.
|
| 15189 |
-
"acc": 0.
|
| 15190 |
-
"ece": 0.
|
| 15191 |
-
"brier": 0.
|
| 15192 |
-
"mean_conf": 0.
|
| 15193 |
-
"confident_error_rate": 0.
|
| 15194 |
-
"coverage_at_0_9": 0.
|
| 15195 |
-
"accuracy_at_0_9": 0.
|
| 15196 |
-
"coverage_at_5pct_error": 0.
|
| 15197 |
-
"coverage_at_1pct_error": 0.
|
| 15198 |
-
"aurc": 0.
|
| 15199 |
-
"error_rate_at_0_9": 0.
|
| 15200 |
-
"confidence_bias": 0.
|
| 15201 |
-
"top_bins": {
|
| 15202 |
-
"0.9": {
|
| 15203 |
-
"n":
|
| 15204 |
-
"errors":
|
| 15205 |
-
"error_rate": 0.
|
| 15206 |
-
},
|
| 15207 |
-
"0.95": {
|
| 15208 |
-
"n":
|
| 15209 |
-
"errors":
|
| 15210 |
-
"error_rate": 0.
|
| 15211 |
},
|
| 15212 |
"0.99": {
|
| 15213 |
-
"n":
|
| 15214 |
-
"errors":
|
| 15215 |
-
"error_rate": 0.
|
| 15216 |
}
|
| 15217 |
},
|
| 15218 |
"selective": {
|
| 15219 |
"0.5": {
|
| 15220 |
"coverage": 0.5,
|
| 15221 |
-
"accuracy": 0.
|
| 15222 |
-
"confidence_cutoff": 0.
|
| 15223 |
},
|
| 15224 |
"0.8": {
|
| 15225 |
-
"coverage": 0.
|
| 15226 |
-
"accuracy": 0.
|
| 15227 |
-
"confidence_cutoff": 0.
|
| 15228 |
}
|
| 15229 |
},
|
| 15230 |
"score_mae": 0.49813993786469746,
|
|
@@ -15232,19 +15232,19 @@
|
|
| 15232 |
},
|
| 15233 |
"none_absent": {
|
| 15234 |
"n": 36,
|
| 15235 |
-
"nll":
|
| 15236 |
"acc": 0.2777777777777778,
|
| 15237 |
-
"ece": 0.
|
| 15238 |
-
"brier":
|
| 15239 |
-
"mean_conf": 0.
|
| 15240 |
"confident_error_rate": 0.05555555555555555,
|
| 15241 |
"coverage_at_0_9": 0.1111111111111111,
|
| 15242 |
"accuracy_at_0_9": 0.5,
|
| 15243 |
"coverage_at_5pct_error": 0.0,
|
| 15244 |
"coverage_at_1pct_error": 0.0,
|
| 15245 |
-
"aurc": 0.
|
| 15246 |
"error_rate_at_0_9": 0.5,
|
| 15247 |
-
"confidence_bias": 0.
|
| 15248 |
"top_bins": {
|
| 15249 |
"0.9": {
|
| 15250 |
"n": 4,
|
|
@@ -15265,39 +15265,39 @@
|
|
| 15265 |
"selective": {
|
| 15266 |
"0.5": {
|
| 15267 |
"coverage": 0.5,
|
| 15268 |
-
"accuracy": 0.
|
| 15269 |
-
"confidence_cutoff": 0.
|
| 15270 |
},
|
| 15271 |
"0.8": {
|
| 15272 |
"coverage": 0.8055555555555556,
|
| 15273 |
-
"accuracy": 0.
|
| 15274 |
-
"confidence_cutoff": 0.
|
| 15275 |
}
|
| 15276 |
}
|
| 15277 |
},
|
| 15278 |
"none_present": {
|
| 15279 |
"n": 36,
|
| 15280 |
-
"nll":
|
| 15281 |
-
"acc": 0.
|
| 15282 |
-
"ece": 0.
|
| 15283 |
-
"brier": 0.
|
| 15284 |
-
"mean_conf": 0.
|
| 15285 |
"confident_error_rate": 0.0,
|
| 15286 |
-
"coverage_at_0_9": 0.
|
| 15287 |
"accuracy_at_0_9": 1.0,
|
| 15288 |
-
"coverage_at_5pct_error": 0.
|
| 15289 |
-
"coverage_at_1pct_error": 0.
|
| 15290 |
-
"aurc": 0.
|
| 15291 |
"error_rate_at_0_9": 0.0,
|
| 15292 |
-
"confidence_bias":
|
| 15293 |
"top_bins": {
|
| 15294 |
"0.9": {
|
| 15295 |
-
"n":
|
| 15296 |
"errors": 0,
|
| 15297 |
"error_rate": 0.0
|
| 15298 |
},
|
| 15299 |
"0.95": {
|
| 15300 |
-
"n":
|
| 15301 |
"errors": 0,
|
| 15302 |
"error_rate": 0.0
|
| 15303 |
},
|
|
@@ -15310,39 +15310,39 @@
|
|
| 15310 |
"selective": {
|
| 15311 |
"0.5": {
|
| 15312 |
"coverage": 0.5,
|
| 15313 |
-
"accuracy":
|
| 15314 |
-
"confidence_cutoff": 0.
|
| 15315 |
},
|
| 15316 |
"0.8": {
|
| 15317 |
"coverage": 0.8055555555555556,
|
| 15318 |
-
"accuracy": 0.
|
| 15319 |
-
"confidence_cutoff": 0.
|
| 15320 |
}
|
| 15321 |
}
|
| 15322 |
},
|
| 15323 |
"permuted": {
|
| 15324 |
"n": 36,
|
| 15325 |
-
"nll": 0.
|
| 15326 |
"acc": 0.6666666666666666,
|
| 15327 |
-
"ece": 0.
|
| 15328 |
-
"brier": 0.
|
| 15329 |
-
"mean_conf": 0.
|
| 15330 |
"confident_error_rate": 0.0,
|
| 15331 |
-
"coverage_at_0_9": 0.
|
| 15332 |
"accuracy_at_0_9": 1.0,
|
| 15333 |
-
"coverage_at_5pct_error": 0.
|
| 15334 |
-
"coverage_at_1pct_error": 0.
|
| 15335 |
-
"aurc": 0.
|
| 15336 |
"error_rate_at_0_9": 0.0,
|
| 15337 |
-
"confidence_bias":
|
| 15338 |
"top_bins": {
|
| 15339 |
"0.9": {
|
| 15340 |
-
"n":
|
| 15341 |
"errors": 0,
|
| 15342 |
"error_rate": 0.0
|
| 15343 |
},
|
| 15344 |
"0.95": {
|
| 15345 |
-
"n":
|
| 15346 |
"errors": 0,
|
| 15347 |
"error_rate": 0.0
|
| 15348 |
},
|
|
@@ -15355,13 +15355,13 @@
|
|
| 15355 |
"selective": {
|
| 15356 |
"0.5": {
|
| 15357 |
"coverage": 0.5,
|
| 15358 |
-
"accuracy":
|
| 15359 |
-
"confidence_cutoff": 0.
|
| 15360 |
},
|
| 15361 |
"0.8": {
|
| 15362 |
"coverage": 0.8055555555555556,
|
| 15363 |
-
"accuracy": 0.
|
| 15364 |
-
"confidence_cutoff": 0.
|
| 15365 |
}
|
| 15366 |
}
|
| 15367 |
}
|
|
@@ -15369,52 +15369,52 @@
|
|
| 15369 |
"heldout_tasks": {},
|
| 15370 |
"permutation": {
|
| 15371 |
"n": 36,
|
| 15372 |
-
"mean_max_delta": 0.
|
| 15373 |
-
"flip_rate": 0.
|
| 15374 |
},
|
| 15375 |
"temperature": 1.0,
|
| 15376 |
"calibrated_clean": {
|
| 15377 |
"n": 1046,
|
| 15378 |
-
"nll": 0.
|
| 15379 |
-
"acc": 0.
|
| 15380 |
-
"ece": 0.
|
| 15381 |
-
"brier": 0.
|
| 15382 |
-
"mean_conf": 0.
|
| 15383 |
-
"confident_error_rate": 0.
|
| 15384 |
-
"coverage_at_0_9": 0.
|
| 15385 |
-
"accuracy_at_0_9": 0.
|
| 15386 |
-
"coverage_at_5pct_error": 0.
|
| 15387 |
-
"coverage_at_1pct_error": 0.
|
| 15388 |
-
"aurc": 0.
|
| 15389 |
-
"error_rate_at_0_9": 0.
|
| 15390 |
-
"confidence_bias": 0.
|
| 15391 |
"top_bins": {
|
| 15392 |
"0.9": {
|
| 15393 |
-
"n":
|
| 15394 |
-
"errors":
|
| 15395 |
-
"error_rate": 0.
|
| 15396 |
},
|
| 15397 |
"0.95": {
|
| 15398 |
-
"n":
|
| 15399 |
-
"errors":
|
| 15400 |
-
"error_rate": 0.
|
| 15401 |
},
|
| 15402 |
"0.99": {
|
| 15403 |
-
"n":
|
| 15404 |
-
"errors":
|
| 15405 |
-
"error_rate": 0.
|
| 15406 |
}
|
| 15407 |
},
|
| 15408 |
"selective": {
|
| 15409 |
"0.5": {
|
| 15410 |
"coverage": 0.5,
|
| 15411 |
-
"accuracy": 0.
|
| 15412 |
-
"confidence_cutoff": 0.
|
| 15413 |
},
|
| 15414 |
"0.8": {
|
| 15415 |
"coverage": 0.8001912045889101,
|
| 15416 |
-
"accuracy": 0.
|
| 15417 |
-
"confidence_cutoff": 0.
|
| 15418 |
}
|
| 15419 |
},
|
| 15420 |
"score_mae": 0.35059676214309027,
|
|
@@ -66859,760 +66859,6 @@
|
|
| 66859 |
}
|
| 66860 |
},
|
| 66861 |
"paired_comparisons": {
|
| 66862 |
-
"Nox minus Sol": {
|
| 66863 |
-
"old_core": {
|
| 66864 |
-
"delta": 0.09255952380952381,
|
| 66865 |
-
"exact_fraction": "311/3360",
|
| 66866 |
-
"paired_ci95": [
|
| 66867 |
-
0.055577566964285605,
|
| 66868 |
-
0.1305831473214286
|
| 66869 |
-
]
|
| 66870 |
-
},
|
| 66871 |
-
"v3_core": {
|
| 66872 |
-
"delta": 0.05708333333333333,
|
| 66873 |
-
"exact_fraction": "137/2400",
|
| 66874 |
-
"paired_ci95": [
|
| 66875 |
-
0.0245833333333334,
|
| 66876 |
-
0.08875
|
| 66877 |
-
]
|
| 66878 |
-
},
|
| 66879 |
-
"v4": {
|
| 66880 |
-
"delta": 0.025,
|
| 66881 |
-
"exact_fraction": "1/40",
|
| 66882 |
-
"paired_ci95": [
|
| 66883 |
-
-0.006288598445143156,
|
| 66884 |
-
0.058030063291139244
|
| 66885 |
-
]
|
| 66886 |
-
},
|
| 66887 |
-
"v5": {
|
| 66888 |
-
"delta": 0.020833333333333332,
|
| 66889 |
-
"exact_fraction": "1/48",
|
| 66890 |
-
"paired_ci95": [
|
| 66891 |
-
-0.006398928762210043,
|
| 66892 |
-
0.04862658709958234
|
| 66893 |
-
]
|
| 66894 |
-
},
|
| 66895 |
-
"transfer_v9_test": {
|
| 66896 |
-
"delta": 0.1089866156787763,
|
| 66897 |
-
"exact_fraction": "57/523",
|
| 66898 |
-
"paired_ci95": [
|
| 66899 |
-
0.07972997762299643,
|
| 66900 |
-
0.13779904306220103
|
| 66901 |
-
]
|
| 66902 |
-
},
|
| 66903 |
-
"original_four_panel_mean": {
|
| 66904 |
-
"delta": 0.04886904761904762,
|
| 66905 |
-
"exact_fraction": "821/16800",
|
| 66906 |
-
"paired_ci95": [
|
| 66907 |
-
0.03301257147338517,
|
| 66908 |
-
0.06502318612678312
|
| 66909 |
-
]
|
| 66910 |
-
},
|
| 66911 |
-
"weighted_mean": {
|
| 66912 |
-
"delta": 0.06526168282800691,
|
| 66913 |
-
"exact_fraction": "2293661/35145600",
|
| 66914 |
-
"paired_ci95": [
|
| 66915 |
-
0.04958814518629331,
|
| 66916 |
-
0.08098195894796445
|
| 66917 |
-
]
|
| 66918 |
-
}
|
| 66919 |
-
},
|
| 66920 |
-
"Nox minus kev-0.8b": {
|
| 66921 |
-
"old_core": {
|
| 66922 |
-
"delta": 0.22864583333333333,
|
| 66923 |
-
"exact_fraction": "439/1920",
|
| 66924 |
-
"paired_ci95": [
|
| 66925 |
-
0.17823567708333335,
|
| 66926 |
-
0.2778273809523808
|
| 66927 |
-
]
|
| 66928 |
-
},
|
| 66929 |
-
"v3_core": {
|
| 66930 |
-
"delta": 0.095,
|
| 66931 |
-
"exact_fraction": "19/200",
|
| 66932 |
-
"paired_ci95": [
|
| 66933 |
-
0.0633333333333333,
|
| 66934 |
-
0.12791666666666673
|
| 66935 |
-
]
|
| 66936 |
-
},
|
| 66937 |
-
"v4": {
|
| 66938 |
-
"delta": 0.1125,
|
| 66939 |
-
"exact_fraction": "9/80",
|
| 66940 |
-
"paired_ci95": [
|
| 66941 |
-
0.06741352201257866,
|
| 66942 |
-
0.15727848715504358
|
| 66943 |
-
]
|
| 66944 |
-
},
|
| 66945 |
-
"v5": {
|
| 66946 |
-
"delta": 0.175,
|
| 66947 |
-
"exact_fraction": "7/40",
|
| 66948 |
-
"paired_ci95": [
|
| 66949 |
-
0.1355393511083524,
|
| 66950 |
-
0.21456521035365878
|
| 66951 |
-
]
|
| 66952 |
-
},
|
| 66953 |
-
"transfer_v9_test": {
|
| 66954 |
-
"delta": 0.06787762906309751,
|
| 66955 |
-
"exact_fraction": "71/1046",
|
| 66956 |
-
"paired_ci95": [
|
| 66957 |
-
0.03431758137001224,
|
| 66958 |
-
0.10142248780822852
|
| 66959 |
-
]
|
| 66960 |
-
},
|
| 66961 |
-
"original_four_panel_mean": {
|
| 66962 |
-
"delta": 0.15278645833333335,
|
| 66963 |
-
"exact_fraction": "5867/38400",
|
| 66964 |
-
"paired_ci95": [
|
| 66965 |
-
0.13186326764793416,
|
| 66966 |
-
0.1736005686577057
|
| 66967 |
-
]
|
| 66968 |
-
},
|
| 66969 |
-
"weighted_mean": {
|
| 66970 |
-
"delta": 0.1456503943594646,
|
| 66971 |
-
"exact_fraction": "487521/3347200",
|
| 66972 |
-
"paired_ci95": [
|
| 66973 |
-
0.12568410699393018,
|
| 66974 |
-
0.16551635549679958
|
| 66975 |
-
]
|
| 66976 |
-
}
|
| 66977 |
-
},
|
| 66978 |
-
"Nox minus kev-4b": {
|
| 66979 |
-
"old_core": {
|
| 66980 |
-
"delta": 0.11097470238095238,
|
| 66981 |
-
"exact_fraction": "2983/26880",
|
| 66982 |
-
"paired_ci95": [
|
| 66983 |
-
0.068154761904762,
|
| 66984 |
-
0.1539462425595238
|
| 66985 |
-
]
|
| 66986 |
-
},
|
| 66987 |
-
"v3_core": {
|
| 66988 |
-
"delta": 0.0325,
|
| 66989 |
-
"exact_fraction": "13/400",
|
| 66990 |
-
"paired_ci95": [
|
| 66991 |
-
0.0033333333333334103,
|
| 66992 |
-
0.06166666666666676
|
| 66993 |
-
]
|
| 66994 |
-
},
|
| 66995 |
-
"v4": {
|
| 66996 |
-
"delta": -0.028125,
|
| 66997 |
-
"exact_fraction": "-9/320",
|
| 66998 |
-
"paired_ci95": [
|
| 66999 |
-
-0.062443225931676956,
|
| 67000 |
-
0.00613059214504551
|
| 67001 |
-
]
|
| 67002 |
-
},
|
| 67003 |
-
"v5": {
|
| 67004 |
-
"delta": 0.016666666666666666,
|
| 67005 |
-
"exact_fraction": "1/60",
|
| 67006 |
-
"paired_ci95": [
|
| 67007 |
-
-0.014652826016486136,
|
| 67008 |
-
0.04855467675421749
|
| 67009 |
-
]
|
| 67010 |
-
},
|
| 67011 |
-
"transfer_v9_test": {
|
| 67012 |
-
"delta": -0.08126195028680688,
|
| 67013 |
-
"exact_fraction": "-85/1046",
|
| 67014 |
-
"paired_ci95": [
|
| 67015 |
-
-0.1091449740378478,
|
| 67016 |
-
-0.05362143534334946
|
| 67017 |
-
]
|
| 67018 |
-
},
|
| 67019 |
-
"original_four_panel_mean": {
|
| 67020 |
-
"delta": 0.03300409226190476,
|
| 67021 |
-
"exact_fraction": "17743/537600",
|
| 67022 |
-
"paired_ci95": [
|
| 67023 |
-
0.0159171480237136,
|
| 67024 |
-
0.05071287650329332
|
| 67025 |
-
]
|
| 67026 |
-
},
|
| 67027 |
-
"weighted_mean": {
|
| 67028 |
-
"delta": 0.027509368171264682,
|
| 67029 |
-
"exact_fraction": "1289111/46860800",
|
| 67030 |
-
"paired_ci95": [
|
| 67031 |
-
0.010595587102736993,
|
| 67032 |
-
0.044668398388024416
|
| 67033 |
-
]
|
| 67034 |
-
}
|
| 67035 |
-
},
|
| 67036 |
-
"Nox minus kev-9b": {
|
| 67037 |
-
"old_core": {
|
| 67038 |
-
"delta": 0.06253720238095238,
|
| 67039 |
-
"exact_fraction": "1681/26880",
|
| 67040 |
-
"paired_ci95": [
|
| 67041 |
-
0.024514508928571384,
|
| 67042 |
-
0.10111700148809517
|
| 67043 |
-
]
|
| 67044 |
-
},
|
| 67045 |
-
"v3_core": {
|
| 67046 |
-
"delta": 0.06041666666666667,
|
| 67047 |
-
"exact_fraction": "29/480",
|
| 67048 |
-
"paired_ci95": [
|
| 67049 |
-
0.02416666666666667,
|
| 67050 |
-
0.09708333333333341
|
| 67051 |
-
]
|
| 67052 |
-
},
|
| 67053 |
-
"v4": {
|
| 67054 |
-
"delta": -0.0765625,
|
| 67055 |
-
"exact_fraction": "-49/640",
|
| 67056 |
-
"paired_ci95": [
|
| 67057 |
-
-0.11406273964723933,
|
| 67058 |
-
-0.040840863146287654
|
| 67059 |
-
]
|
| 67060 |
-
},
|
| 67061 |
-
"v5": {
|
| 67062 |
-
"delta": 0.027083333333333334,
|
| 67063 |
-
"exact_fraction": "13/480",
|
| 67064 |
-
"paired_ci95": [
|
| 67065 |
-
-0.008081049575002395,
|
| 67066 |
-
0.06178192588510811
|
| 67067 |
-
]
|
| 67068 |
-
},
|
| 67069 |
-
"transfer_v9_test": {
|
| 67070 |
-
"delta": -0.11281070745697896,
|
| 67071 |
-
"exact_fraction": "-59/523",
|
| 67072 |
-
"paired_ci95": [
|
| 67073 |
-
-0.14365179025963015,
|
| 67074 |
-
-0.08235239701801159
|
| 67075 |
-
]
|
| 67076 |
-
},
|
| 67077 |
-
"original_four_panel_mean": {
|
| 67078 |
-
"delta": 0.018368675595238096,
|
| 67079 |
-
"exact_fraction": "395/21504",
|
| 67080 |
-
"paired_ci95": [
|
| 67081 |
-
0.00015929724577673504,
|
| 67082 |
-
0.03635734921356825
|
| 67083 |
-
]
|
| 67084 |
-
},
|
| 67085 |
-
"weighted_mean": {
|
| 67086 |
-
"delta": 0.009521846262405535,
|
| 67087 |
-
"exact_fraction": "334651/35145600",
|
| 67088 |
-
"paired_ci95": [
|
| 67089 |
-
-0.007595435407328911,
|
| 67090 |
-
0.026801027422374526
|
| 67091 |
-
]
|
| 67092 |
-
}
|
| 67093 |
-
},
|
| 67094 |
-
"Nox minus Jev": {
|
| 67095 |
-
"old_core": {
|
| 67096 |
-
"delta": 0.0390625,
|
| 67097 |
-
"exact_fraction": "5/128",
|
| 67098 |
-
"paired_ci95": [
|
| 67099 |
-
0.008703497023809521,
|
| 67100 |
-
0.06897321428571423
|
| 67101 |
-
]
|
| 67102 |
-
},
|
| 67103 |
-
"v3_core": {
|
| 67104 |
-
"delta": -0.14583333333333334,
|
| 67105 |
-
"exact_fraction": "-7/48",
|
| 67106 |
-
"paired_ci95": [
|
| 67107 |
-
-0.18834375000000017,
|
| 67108 |
-
-0.10373958333333338
|
| 67109 |
-
]
|
| 67110 |
-
},
|
| 67111 |
-
"v4": {
|
| 67112 |
-
"delta": -0.1546875,
|
| 67113 |
-
"exact_fraction": "-99/640",
|
| 67114 |
-
"paired_ci95": [
|
| 67115 |
-
-0.19683761503067482,
|
| 67116 |
-
-0.11406250000000007
|
| 67117 |
-
]
|
| 67118 |
-
},
|
| 67119 |
-
"v5": {
|
| 67120 |
-
"delta": -0.035416666666666666,
|
| 67121 |
-
"exact_fraction": "-17/480",
|
| 67122 |
-
"paired_ci95": [
|
| 67123 |
-
-0.06684327822543287,
|
| 67124 |
-
-0.0039144015898180265
|
| 67125 |
-
]
|
| 67126 |
-
},
|
| 67127 |
-
"transfer_v9_test": {
|
| 67128 |
-
"delta": -0.1921606118546845,
|
| 67129 |
-
"exact_fraction": "-201/1046",
|
| 67130 |
-
"paired_ci95": [
|
| 67131 |
-
-0.2229480709592455,
|
| 67132 |
-
-0.16162570888468808
|
| 67133 |
-
]
|
| 67134 |
-
},
|
| 67135 |
-
"original_four_panel_mean": {
|
| 67136 |
-
"delta": -0.07421875,
|
| 67137 |
-
"exact_fraction": "-19/256",
|
| 67138 |
-
"paired_ci95": [
|
| 67139 |
-
-0.09309387400474764,
|
| 67140 |
-
-0.05577344994758131
|
| 67141 |
-
]
|
| 67142 |
-
},
|
| 67143 |
-
"weighted_mean": {
|
| 67144 |
-
"delta": -0.08207930011153601,
|
| 67145 |
-
"exact_fraction": "-329683/4016640",
|
| 67146 |
-
"paired_ci95": [
|
| 67147 |
-
-0.09913724619038332,
|
| 67148 |
-
-0.06522481396228826
|
| 67149 |
-
]
|
| 67150 |
-
}
|
| 67151 |
-
},
|
| 67152 |
-
"Nox minus Laya-base": {
|
| 67153 |
-
"old_core": {
|
| 67154 |
-
"delta": 0.26458333333333334,
|
| 67155 |
-
"exact_fraction": "127/480",
|
| 67156 |
-
"paired_ci95": [
|
| 67157 |
-
0.21428478422619038,
|
| 67158 |
-
0.31428664434523806
|
| 67159 |
-
]
|
| 67160 |
-
},
|
| 67161 |
-
"v3_core": {
|
| 67162 |
-
"delta": 0.16458333333333333,
|
| 67163 |
-
"exact_fraction": "79/480",
|
| 67164 |
-
"paired_ci95": [
|
| 67165 |
-
0.12958333333333344,
|
| 67166 |
-
0.19874999999999993
|
| 67167 |
-
]
|
| 67168 |
-
},
|
| 67169 |
-
"v4": {
|
| 67170 |
-
"delta": 0.2765625,
|
| 67171 |
-
"exact_fraction": "177/640",
|
| 67172 |
-
"paired_ci95": [
|
| 67173 |
-
0.2216732154792797,
|
| 67174 |
-
0.3312538580246914
|
| 67175 |
-
]
|
| 67176 |
-
},
|
| 67177 |
-
"v5": {
|
| 67178 |
-
"delta": 0.225,
|
| 67179 |
-
"exact_fraction": "9/40",
|
| 67180 |
-
"paired_ci95": [
|
| 67181 |
-
0.18326279152965536,
|
| 67182 |
-
0.26607286866359453
|
| 67183 |
-
]
|
| 67184 |
-
},
|
| 67185 |
-
"transfer_v9_test": {
|
| 67186 |
-
"delta": 0.1491395793499044,
|
| 67187 |
-
"exact_fraction": "78/523",
|
| 67188 |
-
"paired_ci95": [
|
| 67189 |
-
0.1103515625,
|
| 67190 |
-
0.18714025122546793
|
| 67191 |
-
]
|
| 67192 |
-
},
|
| 67193 |
-
"original_four_panel_mean": {
|
| 67194 |
-
"delta": 0.23268229166666668,
|
| 67195 |
-
"exact_fraction": "1787/7680",
|
| 67196 |
-
"paired_ci95": [
|
| 67197 |
-
0.20941479053264894,
|
| 67198 |
-
0.2554643867871834
|
| 67199 |
-
]
|
| 67200 |
-
},
|
| 67201 |
-
"weighted_mean": {
|
| 67202 |
-
"delta": 0.218126145235819,
|
| 67203 |
-
"exact_fraction": "4380671/20083200",
|
| 67204 |
-
"paired_ci95": [
|
| 67205 |
-
0.19712453859128337,
|
| 67206 |
-
0.23895587623050776
|
| 67207 |
-
]
|
| 67208 |
-
}
|
| 67209 |
-
},
|
| 67210 |
-
"Nox minus Laya-multilingual": {
|
| 67211 |
-
"old_core": {
|
| 67212 |
-
"delta": 0.35751488095238093,
|
| 67213 |
-
"exact_fraction": "961/2688",
|
| 67214 |
-
"paired_ci95": [
|
| 67215 |
-
0.3083305431547619,
|
| 67216 |
-
0.40416666666666656
|
| 67217 |
-
]
|
| 67218 |
-
},
|
| 67219 |
-
"v3_core": {
|
| 67220 |
-
"delta": 0.12875,
|
| 67221 |
-
"exact_fraction": "103/800",
|
| 67222 |
-
"paired_ci95": [
|
| 67223 |
-
0.09583333333333333,
|
| 67224 |
-
0.16166666666666668
|
| 67225 |
-
]
|
| 67226 |
-
},
|
| 67227 |
-
"v4": {
|
| 67228 |
-
"delta": 0.2828125,
|
| 67229 |
-
"exact_fraction": "181/640",
|
| 67230 |
-
"paired_ci95": [
|
| 67231 |
-
0.2321340794423215,
|
| 67232 |
-
0.33590875589622643
|
| 67233 |
-
]
|
| 67234 |
-
},
|
| 67235 |
-
"v5": {
|
| 67236 |
-
"delta": 0.28958333333333336,
|
| 67237 |
-
"exact_fraction": "139/480",
|
| 67238 |
-
"paired_ci95": [
|
| 67239 |
-
0.24203304326814726,
|
| 67240 |
-
0.33659227150537624
|
| 67241 |
-
]
|
| 67242 |
-
},
|
| 67243 |
-
"transfer_v9_test": {
|
| 67244 |
-
"delta": 0.2084130019120459,
|
| 67245 |
-
"exact_fraction": "109/523",
|
| 67246 |
-
"paired_ci95": [
|
| 67247 |
-
0.16923040152963667,
|
| 67248 |
-
0.2462693477059149
|
| 67249 |
-
]
|
| 67250 |
-
},
|
| 67251 |
-
"original_four_panel_mean": {
|
| 67252 |
-
"delta": 0.26466517857142857,
|
| 67253 |
-
"exact_fraction": "11857/44800",
|
| 67254 |
-
"paired_ci95": [
|
| 67255 |
-
0.24161685752065346,
|
| 67256 |
-
0.2875022282879109
|
| 67257 |
-
]
|
| 67258 |
-
},
|
| 67259 |
-
"weighted_mean": {
|
| 67260 |
-
"delta": 0.25656328957252117,
|
| 67261 |
-
"exact_fraction": "12022761/46860800",
|
| 67262 |
-
"paired_ci95": [
|
| 67263 |
-
0.23600300131880433,
|
| 67264 |
-
0.27696035962801896
|
| 67265 |
-
]
|
| 67266 |
-
}
|
| 67267 |
-
},
|
| 67268 |
-
"Nox minus Qwen3.5-9B": {
|
| 67269 |
-
"old_core": {
|
| 67270 |
-
"delta": 0.09088541666666666,
|
| 67271 |
-
"exact_fraction": "349/3840",
|
| 67272 |
-
"paired_ci95": [
|
| 67273 |
-
0.05066964285714293,
|
| 67274 |
-
0.13192057291666676
|
| 67275 |
-
]
|
| 67276 |
-
},
|
| 67277 |
-
"v3_core": {
|
| 67278 |
-
"delta": 0.07166666666666667,
|
| 67279 |
-
"exact_fraction": "43/600",
|
| 67280 |
-
"paired_ci95": [
|
| 67281 |
-
0.038333333333333275,
|
| 67282 |
-
0.10499999999999993
|
| 67283 |
-
]
|
| 67284 |
-
},
|
| 67285 |
-
"v4": {
|
| 67286 |
-
"delta": -0.1078125,
|
| 67287 |
-
"exact_fraction": "-69/640",
|
| 67288 |
-
"paired_ci95": [
|
| 67289 |
-
-0.15406680969159642,
|
| 67290 |
-
-0.0620528650542495
|
| 67291 |
-
]
|
| 67292 |
-
},
|
| 67293 |
-
"v5": {
|
| 67294 |
-
"delta": 0.06666666666666667,
|
| 67295 |
-
"exact_fraction": "1/15",
|
| 67296 |
-
"paired_ci95": [
|
| 67297 |
-
0.027479948547215406,
|
| 67298 |
-
0.10597745181290595
|
| 67299 |
-
]
|
| 67300 |
-
},
|
| 67301 |
-
"transfer_v9_test": {
|
| 67302 |
-
"delta": -0.052581261950286805,
|
| 67303 |
-
"exact_fraction": "-55/1046",
|
| 67304 |
-
"paired_ci95": [
|
| 67305 |
-
-0.08596353013081821,
|
| 67306 |
-
-0.01888529584691864
|
| 67307 |
-
]
|
| 67308 |
-
},
|
| 67309 |
-
"original_four_panel_mean": {
|
| 67310 |
-
"delta": 0.0303515625,
|
| 67311 |
-
"exact_fraction": "777/25600",
|
| 67312 |
-
"paired_ci95": [
|
| 67313 |
-
0.01038540473033447,
|
| 67314 |
-
0.050420589249771885
|
| 67315 |
-
]
|
| 67316 |
-
},
|
| 67317 |
-
"weighted_mean": {
|
| 67318 |
-
"delta": 0.031123227374123645,
|
| 67319 |
-
"exact_fraction": "312527/10041600",
|
| 67320 |
-
"paired_ci95": [
|
| 67321 |
-
0.013295818959364762,
|
| 67322 |
-
0.04931011228775257
|
| 67323 |
-
]
|
| 67324 |
-
}
|
| 67325 |
-
},
|
| 67326 |
-
"Nox minus Lux": {
|
| 67327 |
-
"old_core": {
|
| 67328 |
-
"delta": -0.0020833333333333333,
|
| 67329 |
-
"exact_fraction": "-1/480",
|
| 67330 |
-
"paired_ci95": [
|
| 67331 |
-
-0.03712797619047614,
|
| 67332 |
-
0.03303664434523808
|
| 67333 |
-
]
|
| 67334 |
-
},
|
| 67335 |
-
"v3_core": {
|
| 67336 |
-
"delta": -0.0008333333333333334,
|
| 67337 |
-
"exact_fraction": "-1/1200",
|
| 67338 |
-
"paired_ci95": [
|
| 67339 |
-
-0.0462499999999999,
|
| 67340 |
-
0.043749999999999956
|
| 67341 |
-
]
|
| 67342 |
-
},
|
| 67343 |
-
"v4": {
|
| 67344 |
-
"delta": -0.1125,
|
| 67345 |
-
"exact_fraction": "-9/80",
|
| 67346 |
-
"paired_ci95": [
|
| 67347 |
-
-0.15226071147798748,
|
| 67348 |
-
-0.07448441096087467
|
| 67349 |
-
]
|
| 67350 |
-
},
|
| 67351 |
-
"v5": {
|
| 67352 |
-
"delta": -0.04583333333333333,
|
| 67353 |
-
"exact_fraction": "-11/240",
|
| 67354 |
-
"paired_ci95": [
|
| 67355 |
-
-0.07416056922725012,
|
| 67356 |
-
-0.018749569559228588
|
| 67357 |
-
]
|
| 67358 |
-
},
|
| 67359 |
-
"transfer_v9_test": {
|
| 67360 |
-
"delta": -0.09464627151051626,
|
| 67361 |
-
"exact_fraction": "-99/1046",
|
| 67362 |
-
"paired_ci95": [
|
| 67363 |
-
-0.12101514862804885,
|
| 67364 |
-
-0.06889873160603051
|
| 67365 |
-
]
|
| 67366 |
-
},
|
| 67367 |
-
"original_four_panel_mean": {
|
| 67368 |
-
"delta": -0.0403125,
|
| 67369 |
-
"exact_fraction": "-129/3200",
|
| 67370 |
-
"paired_ci95": [
|
| 67371 |
-
-0.05905983411444227,
|
| 67372 |
-
-0.02179353585296119
|
| 67373 |
-
]
|
| 67374 |
-
},
|
| 67375 |
-
"weighted_mean": {
|
| 67376 |
-
"delta": -0.038780274059910774,
|
| 67377 |
-
"exact_fraction": "-48677/1255200",
|
| 67378 |
-
"paired_ci95": [
|
| 67379 |
-
-0.05625030602618532,
|
| 67380 |
-
-0.021731519294624788
|
| 67381 |
-
]
|
| 67382 |
-
}
|
| 67383 |
-
},
|
| 67384 |
-
"Nox minus Decider": {
|
| 67385 |
-
"old_core": {
|
| 67386 |
-
"delta": 0.18988095238095237,
|
| 67387 |
-
"exact_fraction": "319/1680",
|
| 67388 |
-
"paired_ci95": [
|
| 67389 |
-
0.1443443080357141,
|
| 67390 |
-
0.23571428571428554
|
| 67391 |
-
]
|
| 67392 |
-
},
|
| 67393 |
-
"v3_core": {
|
| 67394 |
-
"delta": 0.052083333333333336,
|
| 67395 |
-
"exact_fraction": "5/96",
|
| 67396 |
-
"paired_ci95": [
|
| 67397 |
-
0.013333333333333308,
|
| 67398 |
-
0.08999999999999997
|
| 67399 |
-
]
|
| 67400 |
-
},
|
| 67401 |
-
"v4": {
|
| 67402 |
-
"delta": -0.1296875,
|
| 67403 |
-
"exact_fraction": "-83/640",
|
| 67404 |
-
"paired_ci95": [
|
| 67405 |
-
-0.17179976851851855,
|
| 67406 |
-
-0.08906249999999993
|
| 67407 |
-
]
|
| 67408 |
-
},
|
| 67409 |
-
"v5": {
|
| 67410 |
-
"delta": 0.01875,
|
| 67411 |
-
"exact_fraction": "3/160",
|
| 67412 |
-
"paired_ci95": [
|
| 67413 |
-
-0.012535333369549428,
|
| 67414 |
-
0.04969926075268807
|
| 67415 |
-
]
|
| 67416 |
-
},
|
| 67417 |
-
"transfer_v9_test": {
|
| 67418 |
-
"delta": -0.01338432122370937,
|
| 67419 |
-
"exact_fraction": "-7/523",
|
| 67420 |
-
"paired_ci95": [
|
| 67421 |
-
-0.0490208104831727,
|
| 67422 |
-
0.022202716340089176
|
| 67423 |
-
]
|
| 67424 |
-
},
|
| 67425 |
-
"original_four_panel_mean": {
|
| 67426 |
-
"delta": 0.03275669642857143,
|
| 67427 |
-
"exact_fraction": "587/17920",
|
| 67428 |
-
"paired_ci95": [
|
| 67429 |
-
0.013155496386521998,
|
| 67430 |
-
0.05228677715137045
|
| 67431 |
-
]
|
| 67432 |
-
},
|
| 67433 |
-
"weighted_mean": {
|
| 67434 |
-
"delta": 0.05133684586406264,
|
| 67435 |
-
"exact_fraction": "7217057/140582400",
|
| 67436 |
-
"paired_ci95": [
|
| 67437 |
-
0.032083278884638945,
|
| 67438 |
-
0.07037086432355429
|
| 67439 |
-
]
|
| 67440 |
-
}
|
| 67441 |
-
},
|
| 67442 |
-
"Nox minus Qwen3.5-2B": {
|
| 67443 |
-
"old_core": {
|
| 67444 |
-
"delta": 0.25877976190476193,
|
| 67445 |
-
"exact_fraction": "1739/6720",
|
| 67446 |
-
"paired_ci95": [
|
| 67447 |
-
0.21566220238095246,
|
| 67448 |
-
0.30174851190476193
|
| 67449 |
-
]
|
| 67450 |
-
},
|
| 67451 |
-
"v3_core": {
|
| 67452 |
-
"delta": 0.12791666666666668,
|
| 67453 |
-
"exact_fraction": "307/2400",
|
| 67454 |
-
"paired_ci95": [
|
| 67455 |
-
0.0979166666666666,
|
| 67456 |
-
0.15999999999999992
|
| 67457 |
-
]
|
| 67458 |
-
},
|
| 67459 |
-
"v4": {
|
| 67460 |
-
"delta": 0.053125,
|
| 67461 |
-
"exact_fraction": "17/320",
|
| 67462 |
-
"paired_ci95": [
|
| 67463 |
-
0.0011595366002794232,
|
| 67464 |
-
0.10381534584912663
|
| 67465 |
-
]
|
| 67466 |
-
},
|
| 67467 |
-
"v5": {
|
| 67468 |
-
"delta": 0.13958333333333334,
|
| 67469 |
-
"exact_fraction": "67/480",
|
| 67470 |
-
"paired_ci95": [
|
| 67471 |
-
0.09772693403789712,
|
| 67472 |
-
0.18099321705426352
|
| 67473 |
-
]
|
| 67474 |
-
},
|
| 67475 |
-
"transfer_v9_test": {
|
| 67476 |
-
"delta": 0.11663479923518165,
|
| 67477 |
-
"exact_fraction": "61/523",
|
| 67478 |
-
"paired_ci95": [
|
| 67479 |
-
0.08030488916814263,
|
| 67480 |
-
0.15300807906502473
|
| 67481 |
-
]
|
| 67482 |
-
},
|
| 67483 |
-
"original_four_panel_mean": {
|
| 67484 |
-
"delta": 0.14485119047619047,
|
| 67485 |
-
"exact_fraction": "4867/33600",
|
| 67486 |
-
"paired_ci95": [
|
| 67487 |
-
0.12338022237889658,
|
| 67488 |
-
0.1658467989420549
|
| 67489 |
-
]
|
| 67490 |
-
},
|
| 67491 |
-
"weighted_mean": {
|
| 67492 |
-
"delta": 0.15601456512337247,
|
| 67493 |
-
"exact_fraction": "10966451/70291200",
|
| 67494 |
-
"paired_ci95": [
|
| 67495 |
-
0.13723175354986636,
|
| 67496 |
-
0.17471297097636712
|
| 67497 |
-
]
|
| 67498 |
-
}
|
| 67499 |
-
},
|
| 67500 |
-
"Nox minus Qwen3.5-4B": {
|
| 67501 |
-
"old_core": {
|
| 67502 |
-
"delta": 0.13113839285714285,
|
| 67503 |
-
"exact_fraction": "235/1792",
|
| 67504 |
-
"paired_ci95": [
|
| 67505 |
-
0.09270647321428567,
|
| 67506 |
-
0.16990420386904764
|
| 67507 |
-
]
|
| 67508 |
-
},
|
| 67509 |
-
"v3_core": {
|
| 67510 |
-
"delta": 0.08458333333333333,
|
| 67511 |
-
"exact_fraction": "203/2400",
|
| 67512 |
-
"paired_ci95": [
|
| 67513 |
-
0.05416666666666664,
|
| 67514 |
-
0.11541666666666661
|
| 67515 |
-
]
|
| 67516 |
-
},
|
| 67517 |
-
"v4": {
|
| 67518 |
-
"delta": -0.0890625,
|
| 67519 |
-
"exact_fraction": "-57/640",
|
| 67520 |
-
"paired_ci95": [
|
| 67521 |
-
-0.1342959349117049,
|
| 67522 |
-
-0.045753734276729595
|
| 67523 |
-
]
|
| 67524 |
-
},
|
| 67525 |
-
"v5": {
|
| 67526 |
-
"delta": 0.06458333333333334,
|
| 67527 |
-
"exact_fraction": "31/480",
|
| 67528 |
-
"paired_ci95": [
|
| 67529 |
-
0.028211805555555566,
|
| 67530 |
-
0.10191039756413821
|
| 67531 |
-
]
|
| 67532 |
-
},
|
| 67533 |
-
"transfer_v9_test": {
|
| 67534 |
-
"delta": -0.008604206500956023,
|
| 67535 |
-
"exact_fraction": "-9/1046",
|
| 67536 |
-
"paired_ci95": [
|
| 67537 |
-
-0.040737677035593334,
|
| 67538 |
-
0.022878932316491855
|
| 67539 |
-
]
|
| 67540 |
-
},
|
| 67541 |
-
"original_four_panel_mean": {
|
| 67542 |
-
"delta": 0.04781063988095238,
|
| 67543 |
-
"exact_fraction": "25703/537600",
|
| 67544 |
-
"paired_ci95": [
|
| 67545 |
-
0.028947119327501308,
|
| 67546 |
-
0.06683696978735856
|
| 67547 |
-
]
|
| 67548 |
-
},
|
| 67549 |
-
"weighted_mean": {
|
| 67550 |
-
"delta": 0.05552484521533279,
|
| 67551 |
-
"exact_fraction": "975727/17572800",
|
| 67552 |
-
"paired_ci95": [
|
| 67553 |
-
0.03826384999615495,
|
| 67554 |
-
0.0725083343854517
|
| 67555 |
-
]
|
| 67556 |
-
}
|
| 67557 |
-
},
|
| 67558 |
-
"Sol minus Nox": {
|
| 67559 |
-
"old_core": {
|
| 67560 |
-
"delta": -0.09255952380952381,
|
| 67561 |
-
"exact_fraction": "-311/3360",
|
| 67562 |
-
"paired_ci95": [
|
| 67563 |
-
-0.13058314732142862,
|
| 67564 |
-
-0.055577566964285646
|
| 67565 |
-
]
|
| 67566 |
-
},
|
| 67567 |
-
"v3_core": {
|
| 67568 |
-
"delta": -0.05708333333333333,
|
| 67569 |
-
"exact_fraction": "-137/2400",
|
| 67570 |
-
"paired_ci95": [
|
| 67571 |
-
-0.08875,
|
| 67572 |
-
-0.0245833333333334
|
| 67573 |
-
]
|
| 67574 |
-
},
|
| 67575 |
-
"v4": {
|
| 67576 |
-
"delta": -0.025,
|
| 67577 |
-
"exact_fraction": "-1/40",
|
| 67578 |
-
"paired_ci95": [
|
| 67579 |
-
-0.058030063291139244,
|
| 67580 |
-
0.006288598445143156
|
| 67581 |
-
]
|
| 67582 |
-
},
|
| 67583 |
-
"v5": {
|
| 67584 |
-
"delta": -0.020833333333333332,
|
| 67585 |
-
"exact_fraction": "-1/48",
|
| 67586 |
-
"paired_ci95": [
|
| 67587 |
-
-0.048626587099582376,
|
| 67588 |
-
0.0063989287622100415
|
| 67589 |
-
]
|
| 67590 |
-
},
|
| 67591 |
-
"transfer_v9_test": {
|
| 67592 |
-
"delta": -0.1089866156787763,
|
| 67593 |
-
"exact_fraction": "-57/523",
|
| 67594 |
-
"paired_ci95": [
|
| 67595 |
-
-0.13779904306220103,
|
| 67596 |
-
-0.07972997762299644
|
| 67597 |
-
]
|
| 67598 |
-
},
|
| 67599 |
-
"original_four_panel_mean": {
|
| 67600 |
-
"delta": -0.04886904761904762,
|
| 67601 |
-
"exact_fraction": "-821/16800",
|
| 67602 |
-
"paired_ci95": [
|
| 67603 |
-
-0.06502318612678314,
|
| 67604 |
-
-0.033012571473385176
|
| 67605 |
-
]
|
| 67606 |
-
},
|
| 67607 |
-
"weighted_mean": {
|
| 67608 |
-
"delta": -0.06526168282800691,
|
| 67609 |
-
"exact_fraction": "-2293661/35145600",
|
| 67610 |
-
"paired_ci95": [
|
| 67611 |
-
-0.08098195894796448,
|
| 67612 |
-
-0.049588145186293314
|
| 67613 |
-
]
|
| 67614 |
-
}
|
| 67615 |
-
},
|
| 67616 |
"Sol minus kev-0.8b": {
|
| 67617 |
"old_core": {
|
| 67618 |
"delta": 0.1360863095238095,
|
|
@@ -68251,64 +67497,6 @@
|
|
| 68251 |
]
|
| 68252 |
}
|
| 68253 |
},
|
| 68254 |
-
"Lux minus Nox": {
|
| 68255 |
-
"old_core": {
|
| 68256 |
-
"delta": 0.0020833333333333333,
|
| 68257 |
-
"exact_fraction": "1/480",
|
| 68258 |
-
"paired_ci95": [
|
| 68259 |
-
-0.033036644345238085,
|
| 68260 |
-
0.03712797619047614
|
| 68261 |
-
]
|
| 68262 |
-
},
|
| 68263 |
-
"v3_core": {
|
| 68264 |
-
"delta": 0.0008333333333333334,
|
| 68265 |
-
"exact_fraction": "1/1200",
|
| 68266 |
-
"paired_ci95": [
|
| 68267 |
-
-0.043749999999999956,
|
| 68268 |
-
0.0462499999999999
|
| 68269 |
-
]
|
| 68270 |
-
},
|
| 68271 |
-
"v4": {
|
| 68272 |
-
"delta": 0.1125,
|
| 68273 |
-
"exact_fraction": "9/80",
|
| 68274 |
-
"paired_ci95": [
|
| 68275 |
-
0.07448441096087466,
|
| 68276 |
-
0.15226071147798748
|
| 68277 |
-
]
|
| 68278 |
-
},
|
| 68279 |
-
"v5": {
|
| 68280 |
-
"delta": 0.04583333333333333,
|
| 68281 |
-
"exact_fraction": "11/240",
|
| 68282 |
-
"paired_ci95": [
|
| 68283 |
-
0.018749569559228584,
|
| 68284 |
-
0.07416056922725009
|
| 68285 |
-
]
|
| 68286 |
-
},
|
| 68287 |
-
"transfer_v9_test": {
|
| 68288 |
-
"delta": 0.09464627151051626,
|
| 68289 |
-
"exact_fraction": "99/1046",
|
| 68290 |
-
"paired_ci95": [
|
| 68291 |
-
0.0688987316060305,
|
| 68292 |
-
0.12101514862804882
|
| 68293 |
-
]
|
| 68294 |
-
},
|
| 68295 |
-
"original_four_panel_mean": {
|
| 68296 |
-
"delta": 0.0403125,
|
| 68297 |
-
"exact_fraction": "129/3200",
|
| 68298 |
-
"paired_ci95": [
|
| 68299 |
-
0.02179353585296119,
|
| 68300 |
-
0.05905983411444227
|
| 68301 |
-
]
|
| 68302 |
-
},
|
| 68303 |
-
"weighted_mean": {
|
| 68304 |
-
"delta": 0.038780274059910774,
|
| 68305 |
-
"exact_fraction": "48677/1255200",
|
| 68306 |
-
"paired_ci95": [
|
| 68307 |
-
0.021731519294624784,
|
| 68308 |
-
0.05625030602618531
|
| 68309 |
-
]
|
| 68310 |
-
}
|
| 68311 |
-
},
|
| 68312 |
"Lux minus Sol": {
|
| 68313 |
"old_core": {
|
| 68314 |
"delta": 0.09464285714285714,
|
|
@@ -68946,6 +68134,180 @@
|
|
| 68946 |
0.11044819701768911
|
| 68947 |
]
|
| 68948 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68949 |
}
|
| 68950 |
},
|
| 68951 |
"method": {
|
|
@@ -68973,10 +68335,10 @@
|
|
| 68973 |
"transfer_accuracy": "Upstream ordered probability argmax, clean knowable records, micro average. Refused/missing probabilities count wrong; unknown evidence is not accuracy.",
|
| 68974 |
"transfer_clusters": "Union IDs/group/parent/control/pair plus exact state SHA and ordered-request IDs, built over all variants then scored only clean knowable rows. Complete components jointly resampled; variable sampled denominator. No whole-family grouping.",
|
| 68975 |
"pairing": "Identical sampled source components across all models; five panels sampled independently. Original four marginal CIs preserved.",
|
| 68976 |
-
"accuracy_arithmetic": "Exact rational integer correct/requested
|
| 68977 |
"interval_role": "Descriptive paired95% intervals; no positive-CI release requirement invented.",
|
| 68978 |
"diagnostics": "Full per-task/proper-score/order/unknown-evidence results remain separate; no accuracy/latency/ECE unit mixing.",
|
| 68979 |
-
"exposure": "Observed regression
|
| 68980 |
},
|
| 68981 |
"prior_weighted_scores": {
|
| 68982 |
"Jev": {
|
|
@@ -68996,12 +68358,9 @@
|
|
| 68996 |
]
|
| 68997 |
},
|
| 68998 |
"Nox": {
|
| 68999 |
-
"point": 0.
|
| 69000 |
-
"exact_fraction": "
|
| 69001 |
-
"
|
| 69002 |
-
0.706127064217637,
|
| 69003 |
-
0.7352055352744082
|
| 69004 |
-
]
|
| 69005 |
},
|
| 69006 |
"kev-9b": {
|
| 69007 |
"point": 0.7201455089684057,
|
|
@@ -69086,6 +68445,6 @@
|
|
| 69086 |
},
|
| 69087 |
"hosted_frontier": "All official calls returned jev-1.13.0; closed weight identity/input retention cannot be independently inspected.",
|
| 69088 |
"size_note": "Qwen3.5 family size labels name their official parent checkpoint. Lux deployed text+decision weights total7,940,895,744 parameters; no vision or generativeLMhead in its serving bundle.",
|
| 69089 |
-
"
|
| 69090 |
-
"
|
| 69091 |
}
|
|
|
|
| 9 |
},
|
| 10 |
"metric_design": "User-requested, outcome-informed decision-priority weights; observed regression data, not blind testing.",
|
| 11 |
"protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
|
| 12 |
+
"statistics_sha256": "b319a993ac569ea1fa8da027ccb2195868557e4efe232a5d06091db2884086be",
|
| 13 |
"models": {
|
| 14 |
"Jev": {
|
| 15 |
"panels": {
|
|
|
|
| 10330 |
]
|
| 10331 |
},
|
| 10332 |
"transfer_v9_test": {
|
| 10333 |
+
"point": 0.6959847036328872,
|
| 10334 |
+
"exact_fraction": "364/523",
|
| 10335 |
"ci95": [
|
| 10336 |
+
0.6660212501997199,
|
| 10337 |
+
0.7249533151094627
|
| 10338 |
]
|
| 10339 |
},
|
| 10340 |
"original_four_panel_mean": {
|
|
|
|
| 10346 |
]
|
| 10347 |
},
|
| 10348 |
"weighted_mean": {
|
| 10349 |
+
"point": 0.7308523186401712,
|
| 10350 |
+
"exact_fraction": "102744973/140582400",
|
| 10351 |
"ci95": [
|
| 10352 |
+
0.7157461705500261,
|
| 10353 |
+
0.7455873812383725
|
| 10354 |
]
|
| 10355 |
}
|
| 10356 |
},
|
|
|
|
| 10496 |
"clean": {
|
| 10497 |
"requested": 1046,
|
| 10498 |
"valid_probability_rows": 1046,
|
| 10499 |
+
"correct": 728,
|
| 10500 |
+
"accuracy": 0.6959847036328872,
|
| 10501 |
"full_probability_metrics": {
|
| 10502 |
"n": 1046,
|
| 10503 |
+
"nll": 0.9079076925806586,
|
| 10504 |
+
"acc": 0.6959847036328872,
|
| 10505 |
+
"ece": 0.11803573910263573,
|
| 10506 |
+
"brier": 0.41689550886050714,
|
| 10507 |
+
"mean_conf": 0.814020442735523,
|
| 10508 |
+
"confident_error_rate": 0.058317399617590825,
|
| 10509 |
+
"coverage_at_0_9": 0.5525812619502868,
|
| 10510 |
+
"accuracy_at_0_9": 0.8944636678200693,
|
| 10511 |
+
"coverage_at_5pct_error": 0.36424474187380496,
|
| 10512 |
+
"coverage_at_1pct_error": 0.14340344168260039,
|
| 10513 |
+
"aurc": 0.11088769274123257,
|
| 10514 |
+
"error_rate_at_0_9": 0.10553633217993079,
|
| 10515 |
+
"confidence_bias": 0.11803573910263576,
|
| 10516 |
"top_bins": {
|
| 10517 |
"0.9": {
|
| 10518 |
+
"n": 578,
|
| 10519 |
+
"errors": 61,
|
| 10520 |
+
"error_rate": 0.10553633217993079
|
| 10521 |
},
|
| 10522 |
"0.95": {
|
| 10523 |
+
"n": 491,
|
| 10524 |
+
"errors": 39,
|
| 10525 |
+
"error_rate": 0.07942973523421588
|
| 10526 |
},
|
| 10527 |
"0.99": {
|
| 10528 |
+
"n": 314,
|
| 10529 |
+
"errors": 11,
|
| 10530 |
+
"error_rate": 0.03503184713375796
|
| 10531 |
}
|
| 10532 |
},
|
| 10533 |
"selective": {
|
| 10534 |
"0.5": {
|
| 10535 |
"coverage": 0.5,
|
| 10536 |
+
"accuracy": 0.9101338432122371,
|
| 10537 |
+
"confidence_cutoff": 0.9301925551074537
|
| 10538 |
},
|
| 10539 |
"0.8": {
|
| 10540 |
"coverage": 0.8001912045889101,
|
| 10541 |
+
"accuracy": 0.7921146953405018,
|
| 10542 |
+
"confidence_cutoff": 0.5957019462220806
|
| 10543 |
}
|
| 10544 |
},
|
| 10545 |
"score_mae": 0.35059676214309027,
|
|
|
|
| 10547 |
},
|
| 10548 |
"covered_only_probability_metrics": {
|
| 10549 |
"n": 1046,
|
| 10550 |
+
"nll": 0.9079076925806586,
|
| 10551 |
+
"acc": 0.6959847036328872,
|
| 10552 |
+
"ece": 0.11803573910263573,
|
| 10553 |
+
"brier": 0.41689550886050714,
|
| 10554 |
+
"mean_conf": 0.814020442735523,
|
| 10555 |
+
"confident_error_rate": 0.058317399617590825,
|
| 10556 |
+
"coverage_at_0_9": 0.5525812619502868,
|
| 10557 |
+
"accuracy_at_0_9": 0.8944636678200693,
|
| 10558 |
+
"coverage_at_5pct_error": 0.36424474187380496,
|
| 10559 |
+
"coverage_at_1pct_error": 0.14340344168260039,
|
| 10560 |
+
"aurc": 0.11088769274123257,
|
| 10561 |
+
"error_rate_at_0_9": 0.10553633217993079,
|
| 10562 |
+
"confidence_bias": 0.11803573910263576,
|
| 10563 |
"top_bins": {
|
| 10564 |
"0.9": {
|
| 10565 |
+
"n": 578,
|
| 10566 |
+
"errors": 61,
|
| 10567 |
+
"error_rate": 0.10553633217993079
|
| 10568 |
},
|
| 10569 |
"0.95": {
|
| 10570 |
+
"n": 491,
|
| 10571 |
+
"errors": 39,
|
| 10572 |
+
"error_rate": 0.07942973523421588
|
| 10573 |
},
|
| 10574 |
"0.99": {
|
| 10575 |
+
"n": 314,
|
| 10576 |
+
"errors": 11,
|
| 10577 |
+
"error_rate": 0.03503184713375796
|
| 10578 |
}
|
| 10579 |
},
|
| 10580 |
"selective": {
|
| 10581 |
"0.5": {
|
| 10582 |
"coverage": 0.5,
|
| 10583 |
+
"accuracy": 0.9101338432122371,
|
| 10584 |
+
"confidence_cutoff": 0.9301925551074537
|
| 10585 |
},
|
| 10586 |
"0.8": {
|
| 10587 |
"coverage": 0.8001912045889101,
|
| 10588 |
+
"accuracy": 0.7921146953405018,
|
| 10589 |
+
"confidence_cutoff": 0.5957019462220806
|
| 10590 |
}
|
| 10591 |
},
|
| 10592 |
"score_mae": 0.35059676214309027,
|
|
|
|
| 10598 |
"buried_emotion": {
|
| 10599 |
"requested": 20,
|
| 10600 |
"valid_probability_rows": 20,
|
| 10601 |
+
"correct": 5,
|
| 10602 |
+
"accuracy": 0.25,
|
| 10603 |
"full_probability_metrics": {
|
| 10604 |
"n": 20,
|
| 10605 |
+
"nll": 1.4541223366826688,
|
| 10606 |
+
"acc": 0.25,
|
| 10607 |
+
"ece": 0.33647159228575807,
|
| 10608 |
+
"brier": 0.784234954627584,
|
| 10609 |
+
"mean_conf": 0.5836932166148608,
|
| 10610 |
"confident_error_rate": 0.0,
|
| 10611 |
+
"coverage_at_0_9": 0.05,
|
| 10612 |
+
"accuracy_at_0_9": 1.0,
|
| 10613 |
+
"coverage_at_5pct_error": 0.05,
|
| 10614 |
+
"coverage_at_1pct_error": 0.05,
|
| 10615 |
+
"aurc": 0.5585412761902699,
|
| 10616 |
+
"error_rate_at_0_9": 0.0,
|
| 10617 |
+
"confidence_bias": 0.33369321661486084,
|
| 10618 |
"top_bins": {
|
| 10619 |
"0.9": {
|
| 10620 |
+
"n": 1,
|
| 10621 |
"errors": 0,
|
| 10622 |
+
"error_rate": 0.0
|
| 10623 |
},
|
| 10624 |
"0.95": {
|
| 10625 |
+
"n": 1,
|
| 10626 |
"errors": 0,
|
| 10627 |
+
"error_rate": 0.0
|
| 10628 |
},
|
| 10629 |
"0.99": {
|
| 10630 |
"n": 0,
|
|
|
|
| 10635 |
"selective": {
|
| 10636 |
"0.5": {
|
| 10637 |
"coverage": 0.5,
|
| 10638 |
+
"accuracy": 0.5,
|
| 10639 |
+
"confidence_cutoff": 0.5472205811051023
|
| 10640 |
},
|
| 10641 |
"0.8": {
|
| 10642 |
"coverage": 0.8,
|
| 10643 |
+
"accuracy": 0.3125,
|
| 10644 |
+
"confidence_cutoff": 0.4314323265237658
|
| 10645 |
}
|
| 10646 |
}
|
| 10647 |
},
|
| 10648 |
"covered_only_probability_metrics": {
|
| 10649 |
"n": 20,
|
| 10650 |
+
"nll": 1.4541223366826688,
|
| 10651 |
+
"acc": 0.25,
|
| 10652 |
+
"ece": 0.33647159228575807,
|
| 10653 |
+
"brier": 0.784234954627584,
|
| 10654 |
+
"mean_conf": 0.5836932166148608,
|
| 10655 |
"confident_error_rate": 0.0,
|
| 10656 |
+
"coverage_at_0_9": 0.05,
|
| 10657 |
+
"accuracy_at_0_9": 1.0,
|
| 10658 |
+
"coverage_at_5pct_error": 0.05,
|
| 10659 |
+
"coverage_at_1pct_error": 0.05,
|
| 10660 |
+
"aurc": 0.5585412761902699,
|
| 10661 |
+
"error_rate_at_0_9": 0.0,
|
| 10662 |
+
"confidence_bias": 0.33369321661486084,
|
| 10663 |
"top_bins": {
|
| 10664 |
"0.9": {
|
| 10665 |
+
"n": 1,
|
| 10666 |
"errors": 0,
|
| 10667 |
+
"error_rate": 0.0
|
| 10668 |
},
|
| 10669 |
"0.95": {
|
| 10670 |
+
"n": 1,
|
| 10671 |
"errors": 0,
|
| 10672 |
+
"error_rate": 0.0
|
| 10673 |
},
|
| 10674 |
"0.99": {
|
| 10675 |
"n": 0,
|
|
|
|
| 10680 |
"selective": {
|
| 10681 |
"0.5": {
|
| 10682 |
"coverage": 0.5,
|
| 10683 |
+
"accuracy": 0.5,
|
| 10684 |
+
"confidence_cutoff": 0.5472205811051023
|
| 10685 |
},
|
| 10686 |
"0.8": {
|
| 10687 |
"coverage": 0.8,
|
| 10688 |
+
"accuracy": 0.3125,
|
| 10689 |
+
"confidence_cutoff": 0.4314323265237658
|
| 10690 |
}
|
| 10691 |
}
|
| 10692 |
},
|
|
|
|
| 11475 |
"emotion": {
|
| 11476 |
"requested": 80,
|
| 11477 |
"valid_probability_rows": 80,
|
| 11478 |
+
"correct": 48,
|
| 11479 |
+
"accuracy": 0.6,
|
| 11480 |
"full_probability_metrics": {
|
| 11481 |
"n": 80,
|
| 11482 |
+
"nll": 1.3247643428213003,
|
| 11483 |
+
"acc": 0.6,
|
| 11484 |
+
"ece": 0.15895568320390185,
|
| 11485 |
+
"brier": 0.5466111456851124,
|
| 11486 |
+
"mean_conf": 0.7589556832039019,
|
| 11487 |
+
"confident_error_rate": 0.05,
|
| 11488 |
+
"coverage_at_0_9": 0.3625,
|
| 11489 |
+
"accuracy_at_0_9": 0.8620689655172413,
|
| 11490 |
+
"coverage_at_5pct_error": 0.25,
|
| 11491 |
+
"coverage_at_1pct_error": 0.0375,
|
| 11492 |
+
"aurc": 0.20182699408158983,
|
| 11493 |
+
"error_rate_at_0_9": 0.13793103448275862,
|
| 11494 |
+
"confidence_bias": 0.15895568320390197,
|
| 11495 |
"top_bins": {
|
| 11496 |
"0.9": {
|
| 11497 |
+
"n": 29,
|
| 11498 |
+
"errors": 4,
|
| 11499 |
+
"error_rate": 0.13793103448275862
|
| 11500 |
},
|
| 11501 |
"0.95": {
|
| 11502 |
+
"n": 19,
|
| 11503 |
+
"errors": 1,
|
| 11504 |
+
"error_rate": 0.05263157894736842
|
| 11505 |
},
|
| 11506 |
"0.99": {
|
| 11507 |
+
"n": 11,
|
| 11508 |
+
"errors": 1,
|
| 11509 |
+
"error_rate": 0.09090909090909091
|
| 11510 |
}
|
| 11511 |
},
|
| 11512 |
"selective": {
|
| 11513 |
"0.5": {
|
| 11514 |
"coverage": 0.5,
|
| 11515 |
+
"accuracy": 0.85,
|
| 11516 |
+
"confidence_cutoff": 0.7988864004860878
|
| 11517 |
},
|
| 11518 |
"0.8": {
|
| 11519 |
"coverage": 0.8,
|
| 11520 |
+
"accuracy": 0.671875,
|
| 11521 |
+
"confidence_cutoff": 0.5632915364363126
|
| 11522 |
}
|
| 11523 |
}
|
| 11524 |
},
|
| 11525 |
"covered_only_probability_metrics": {
|
| 11526 |
"n": 80,
|
| 11527 |
+
"nll": 1.3247643428213003,
|
| 11528 |
+
"acc": 0.6,
|
| 11529 |
+
"ece": 0.15895568320390185,
|
| 11530 |
+
"brier": 0.5466111456851124,
|
| 11531 |
+
"mean_conf": 0.7589556832039019,
|
| 11532 |
+
"confident_error_rate": 0.05,
|
| 11533 |
+
"coverage_at_0_9": 0.3625,
|
| 11534 |
+
"accuracy_at_0_9": 0.8620689655172413,
|
| 11535 |
+
"coverage_at_5pct_error": 0.25,
|
| 11536 |
+
"coverage_at_1pct_error": 0.0375,
|
| 11537 |
+
"aurc": 0.20182699408158983,
|
| 11538 |
+
"error_rate_at_0_9": 0.13793103448275862,
|
| 11539 |
+
"confidence_bias": 0.15895568320390197,
|
| 11540 |
"top_bins": {
|
| 11541 |
"0.9": {
|
| 11542 |
+
"n": 29,
|
| 11543 |
+
"errors": 4,
|
| 11544 |
+
"error_rate": 0.13793103448275862
|
| 11545 |
},
|
| 11546 |
"0.95": {
|
| 11547 |
+
"n": 19,
|
| 11548 |
+
"errors": 1,
|
| 11549 |
+
"error_rate": 0.05263157894736842
|
| 11550 |
},
|
| 11551 |
"0.99": {
|
| 11552 |
+
"n": 11,
|
| 11553 |
+
"errors": 1,
|
| 11554 |
+
"error_rate": 0.09090909090909091
|
| 11555 |
}
|
| 11556 |
},
|
| 11557 |
"selective": {
|
| 11558 |
"0.5": {
|
| 11559 |
"coverage": 0.5,
|
| 11560 |
+
"accuracy": 0.85,
|
| 11561 |
+
"confidence_cutoff": 0.7988864004860878
|
| 11562 |
},
|
| 11563 |
"0.8": {
|
| 11564 |
"coverage": 0.8,
|
| 11565 |
+
"accuracy": 0.671875,
|
| 11566 |
+
"confidence_cutoff": 0.5632915364363126
|
| 11567 |
}
|
| 11568 |
}
|
| 11569 |
},
|
|
|
|
| 13243 |
"coverage": 1.0
|
| 13244 |
}
|
| 13245 |
},
|
| 13246 |
+
"clean_task_macro_accuracy": 0.729675925925926,
|
| 13247 |
"permutation": {
|
| 13248 |
"requested_pairs": 36,
|
| 13249 |
"valid_pairs": 36,
|
| 13250 |
+
"both_correct_requested": 0.6388888888888888,
|
| 13251 |
+
"flip_rate": 0.16666666666666666,
|
| 13252 |
+
"covered_only_flip_rate": 0.16666666666666666,
|
| 13253 |
+
"half_l1": 0.07949225370274392
|
| 13254 |
},
|
| 13255 |
"unknowable": {
|
| 13256 |
"requested": 110,
|
|
|
|
| 13378 |
"truncated_questions": 0
|
| 13379 |
},
|
| 13380 |
"complete_unchanged_upstream_report": {
|
| 13381 |
+
"objective": -0.7834679467690011,
|
| 13382 |
"paired_flip": {
|
| 13383 |
"pairs": 64,
|
| 13384 |
"flip_rate": 0.828125,
|
|
|
|
| 13399 |
},
|
| 13400 |
"clean": {
|
| 13401 |
"n": 1046,
|
| 13402 |
+
"nll": 0.9079076925806586,
|
| 13403 |
+
"acc": 0.6959847036328872,
|
| 13404 |
+
"ece": 0.11803573910263573,
|
| 13405 |
+
"brier": 0.41689550886050714,
|
| 13406 |
+
"mean_conf": 0.814020442735523,
|
| 13407 |
+
"confident_error_rate": 0.058317399617590825,
|
| 13408 |
+
"coverage_at_0_9": 0.5525812619502868,
|
| 13409 |
+
"accuracy_at_0_9": 0.8944636678200693,
|
| 13410 |
+
"coverage_at_5pct_error": 0.36424474187380496,
|
| 13411 |
+
"coverage_at_1pct_error": 0.14340344168260039,
|
| 13412 |
+
"aurc": 0.11088769274123257,
|
| 13413 |
+
"error_rate_at_0_9": 0.10553633217993079,
|
| 13414 |
+
"confidence_bias": 0.11803573910263576,
|
| 13415 |
"top_bins": {
|
| 13416 |
"0.9": {
|
| 13417 |
+
"n": 578,
|
| 13418 |
+
"errors": 61,
|
| 13419 |
+
"error_rate": 0.10553633217993079
|
| 13420 |
},
|
| 13421 |
"0.95": {
|
| 13422 |
+
"n": 491,
|
| 13423 |
+
"errors": 39,
|
| 13424 |
+
"error_rate": 0.07942973523421588
|
| 13425 |
},
|
| 13426 |
"0.99": {
|
| 13427 |
+
"n": 314,
|
| 13428 |
+
"errors": 11,
|
| 13429 |
+
"error_rate": 0.03503184713375796
|
| 13430 |
}
|
| 13431 |
},
|
| 13432 |
"selective": {
|
| 13433 |
"0.5": {
|
| 13434 |
"coverage": 0.5,
|
| 13435 |
+
"accuracy": 0.9101338432122371,
|
| 13436 |
+
"confidence_cutoff": 0.9301925551074537
|
| 13437 |
},
|
| 13438 |
"0.8": {
|
| 13439 |
"coverage": 0.8001912045889101,
|
| 13440 |
+
"accuracy": 0.7921146953405018,
|
| 13441 |
+
"confidence_cutoff": 0.5957019462220806
|
| 13442 |
}
|
| 13443 |
},
|
| 13444 |
"score_mae": 0.35059676214309027,
|
|
|
|
| 13447 |
"tasks": {
|
| 13448 |
"buried_emotion": {
|
| 13449 |
"n": 20,
|
| 13450 |
+
"nll": 1.4541223366826688,
|
| 13451 |
+
"acc": 0.25,
|
| 13452 |
+
"ece": 0.33647159228575807,
|
| 13453 |
+
"brier": 0.784234954627584,
|
| 13454 |
+
"mean_conf": 0.5836932166148608,
|
| 13455 |
"confident_error_rate": 0.0,
|
| 13456 |
+
"coverage_at_0_9": 0.05,
|
| 13457 |
+
"accuracy_at_0_9": 1.0,
|
| 13458 |
+
"coverage_at_5pct_error": 0.05,
|
| 13459 |
+
"coverage_at_1pct_error": 0.05,
|
| 13460 |
+
"aurc": 0.5585412761902699,
|
| 13461 |
+
"error_rate_at_0_9": 0.0,
|
| 13462 |
+
"confidence_bias": 0.33369321661486084,
|
| 13463 |
"top_bins": {
|
| 13464 |
"0.9": {
|
| 13465 |
+
"n": 1,
|
| 13466 |
"errors": 0,
|
| 13467 |
+
"error_rate": 0.0
|
| 13468 |
},
|
| 13469 |
"0.95": {
|
| 13470 |
+
"n": 1,
|
| 13471 |
"errors": 0,
|
| 13472 |
+
"error_rate": 0.0
|
| 13473 |
},
|
| 13474 |
"0.99": {
|
| 13475 |
"n": 0,
|
|
|
|
| 13480 |
"selective": {
|
| 13481 |
"0.5": {
|
| 13482 |
"coverage": 0.5,
|
| 13483 |
+
"accuracy": 0.5,
|
| 13484 |
+
"confidence_cutoff": 0.5472205811051023
|
| 13485 |
},
|
| 13486 |
"0.8": {
|
| 13487 |
"coverage": 0.8,
|
| 13488 |
+
"accuracy": 0.3125,
|
| 13489 |
+
"confidence_cutoff": 0.4314323265237658
|
| 13490 |
}
|
| 13491 |
}
|
| 13492 |
},
|
|
|
|
| 13854 |
},
|
| 13855 |
"emotion": {
|
| 13856 |
"n": 80,
|
| 13857 |
+
"nll": 1.3247643428213003,
|
| 13858 |
+
"acc": 0.6,
|
| 13859 |
+
"ece": 0.15895568320390185,
|
| 13860 |
+
"brier": 0.5466111456851124,
|
| 13861 |
+
"mean_conf": 0.7589556832039019,
|
| 13862 |
+
"confident_error_rate": 0.05,
|
| 13863 |
+
"coverage_at_0_9": 0.3625,
|
| 13864 |
+
"accuracy_at_0_9": 0.8620689655172413,
|
| 13865 |
+
"coverage_at_5pct_error": 0.25,
|
| 13866 |
+
"coverage_at_1pct_error": 0.0375,
|
| 13867 |
+
"aurc": 0.20182699408158983,
|
| 13868 |
+
"error_rate_at_0_9": 0.13793103448275862,
|
| 13869 |
+
"confidence_bias": 0.15895568320390197,
|
| 13870 |
"top_bins": {
|
| 13871 |
"0.9": {
|
| 13872 |
+
"n": 29,
|
| 13873 |
+
"errors": 4,
|
| 13874 |
+
"error_rate": 0.13793103448275862
|
| 13875 |
},
|
| 13876 |
"0.95": {
|
| 13877 |
+
"n": 19,
|
| 13878 |
+
"errors": 1,
|
| 13879 |
+
"error_rate": 0.05263157894736842
|
| 13880 |
},
|
| 13881 |
"0.99": {
|
| 13882 |
+
"n": 11,
|
| 13883 |
+
"errors": 1,
|
| 13884 |
+
"error_rate": 0.09090909090909091
|
| 13885 |
}
|
| 13886 |
},
|
| 13887 |
"selective": {
|
| 13888 |
"0.5": {
|
| 13889 |
"coverage": 0.5,
|
| 13890 |
+
"accuracy": 0.85,
|
| 13891 |
+
"confidence_cutoff": 0.7988864004860878
|
| 13892 |
},
|
| 13893 |
"0.8": {
|
| 13894 |
"coverage": 0.8,
|
| 13895 |
+
"accuracy": 0.671875,
|
| 13896 |
+
"confidence_cutoff": 0.5632915364363126
|
| 13897 |
}
|
| 13898 |
}
|
| 13899 |
},
|
|
|
|
| 15185 |
"variants": {
|
| 15186 |
"clean": {
|
| 15187 |
"n": 1156,
|
| 15188 |
+
"nll": 1.0007605953968302,
|
| 15189 |
+
"acc": 0.6704152249134948,
|
| 15190 |
+
"ece": 0.1409850221700258,
|
| 15191 |
+
"brier": 0.45929116387613567,
|
| 15192 |
+
"mean_conf": 0.8114002470835204,
|
| 15193 |
+
"confident_error_rate": 0.06833910034602077,
|
| 15194 |
+
"coverage_at_0_9": 0.5259515570934256,
|
| 15195 |
+
"accuracy_at_0_9": 0.8700657894736842,
|
| 15196 |
+
"coverage_at_5pct_error": 0.20242214532871972,
|
| 15197 |
+
"coverage_at_1pct_error": 0.12975778546712802,
|
| 15198 |
+
"aurc": 0.1369722577685153,
|
| 15199 |
+
"error_rate_at_0_9": 0.1299342105263158,
|
| 15200 |
+
"confidence_bias": 0.14098502217002562,
|
| 15201 |
+
"top_bins": {
|
| 15202 |
+
"0.9": {
|
| 15203 |
+
"n": 608,
|
| 15204 |
+
"errors": 79,
|
| 15205 |
+
"error_rate": 0.1299342105263158
|
| 15206 |
+
},
|
| 15207 |
+
"0.95": {
|
| 15208 |
+
"n": 515,
|
| 15209 |
+
"errors": 52,
|
| 15210 |
+
"error_rate": 0.10097087378640776
|
| 15211 |
},
|
| 15212 |
"0.99": {
|
| 15213 |
+
"n": 328,
|
| 15214 |
+
"errors": 19,
|
| 15215 |
+
"error_rate": 0.057926829268292686
|
| 15216 |
}
|
| 15217 |
},
|
| 15218 |
"selective": {
|
| 15219 |
"0.5": {
|
| 15220 |
"coverage": 0.5,
|
| 15221 |
+
"accuracy": 0.8788927335640139,
|
| 15222 |
+
"confidence_cutoff": 0.9157804353063602
|
| 15223 |
},
|
| 15224 |
"0.8": {
|
| 15225 |
+
"coverage": 0.801038062283737,
|
| 15226 |
+
"accuracy": 0.7591792656587473,
|
| 15227 |
+
"confidence_cutoff": 0.6007340833161877
|
| 15228 |
}
|
| 15229 |
},
|
| 15230 |
"score_mae": 0.49813993786469746,
|
|
|
|
| 15232 |
},
|
| 15233 |
"none_absent": {
|
| 15234 |
"n": 36,
|
| 15235 |
+
"nll": 2.016479899399725,
|
| 15236 |
"acc": 0.2777777777777778,
|
| 15237 |
+
"ece": 0.43639357499087156,
|
| 15238 |
+
"brier": 1.0193378435955023,
|
| 15239 |
+
"mean_conf": 0.7141713527686493,
|
| 15240 |
"confident_error_rate": 0.05555555555555555,
|
| 15241 |
"coverage_at_0_9": 0.1111111111111111,
|
| 15242 |
"accuracy_at_0_9": 0.5,
|
| 15243 |
"coverage_at_5pct_error": 0.0,
|
| 15244 |
"coverage_at_1pct_error": 0.0,
|
| 15245 |
+
"aurc": 0.6622061759903008,
|
| 15246 |
"error_rate_at_0_9": 0.5,
|
| 15247 |
+
"confidence_bias": 0.4363935749908715,
|
| 15248 |
"top_bins": {
|
| 15249 |
"0.9": {
|
| 15250 |
"n": 4,
|
|
|
|
| 15265 |
"selective": {
|
| 15266 |
"0.5": {
|
| 15267 |
"coverage": 0.5,
|
| 15268 |
+
"accuracy": 0.3888888888888889,
|
| 15269 |
+
"confidence_cutoff": 0.7284490118427092
|
| 15270 |
},
|
| 15271 |
"0.8": {
|
| 15272 |
"coverage": 0.8055555555555556,
|
| 15273 |
+
"accuracy": 0.3448275862068966,
|
| 15274 |
+
"confidence_cutoff": 0.5772219386282852
|
| 15275 |
}
|
| 15276 |
}
|
| 15277 |
},
|
| 15278 |
"none_present": {
|
| 15279 |
"n": 36,
|
| 15280 |
+
"nll": 0.9834989953062324,
|
| 15281 |
+
"acc": 0.6666666666666666,
|
| 15282 |
+
"ece": 0.16246789063665687,
|
| 15283 |
+
"brier": 0.42631332906431735,
|
| 15284 |
+
"mean_conf": 0.7475449294183114,
|
| 15285 |
"confident_error_rate": 0.0,
|
| 15286 |
+
"coverage_at_0_9": 0.3888888888888889,
|
| 15287 |
"accuracy_at_0_9": 1.0,
|
| 15288 |
+
"coverage_at_5pct_error": 0.5,
|
| 15289 |
+
"coverage_at_1pct_error": 0.5,
|
| 15290 |
+
"aurc": 0.10954859179448669,
|
| 15291 |
"error_rate_at_0_9": 0.0,
|
| 15292 |
+
"confidence_bias": 0.08087826275164478,
|
| 15293 |
"top_bins": {
|
| 15294 |
"0.9": {
|
| 15295 |
+
"n": 14,
|
| 15296 |
"errors": 0,
|
| 15297 |
"error_rate": 0.0
|
| 15298 |
},
|
| 15299 |
"0.95": {
|
| 15300 |
+
"n": 12,
|
| 15301 |
"errors": 0,
|
| 15302 |
"error_rate": 0.0
|
| 15303 |
},
|
|
|
|
| 15310 |
"selective": {
|
| 15311 |
"0.5": {
|
| 15312 |
"coverage": 0.5,
|
| 15313 |
+
"accuracy": 1.0,
|
| 15314 |
+
"confidence_cutoff": 0.8038808995404794
|
| 15315 |
},
|
| 15316 |
"0.8": {
|
| 15317 |
"coverage": 0.8055555555555556,
|
| 15318 |
+
"accuracy": 0.7586206896551724,
|
| 15319 |
+
"confidence_cutoff": 0.514763438695478
|
| 15320 |
}
|
| 15321 |
}
|
| 15322 |
},
|
| 15323 |
"permuted": {
|
| 15324 |
"n": 36,
|
| 15325 |
+
"nll": 0.9187906169822195,
|
| 15326 |
"acc": 0.6666666666666666,
|
| 15327 |
+
"ece": 0.13776171917514016,
|
| 15328 |
+
"brier": 0.4066078577529032,
|
| 15329 |
+
"mean_conf": 0.7543481885836296,
|
| 15330 |
"confident_error_rate": 0.0,
|
| 15331 |
+
"coverage_at_0_9": 0.4166666666666667,
|
| 15332 |
"accuracy_at_0_9": 1.0,
|
| 15333 |
+
"coverage_at_5pct_error": 0.5555555555555556,
|
| 15334 |
+
"coverage_at_1pct_error": 0.5,
|
| 15335 |
+
"aurc": 0.09826261699381092,
|
| 15336 |
"error_rate_at_0_9": 0.0,
|
| 15337 |
+
"confidence_bias": 0.08768152191696299,
|
| 15338 |
"top_bins": {
|
| 15339 |
"0.9": {
|
| 15340 |
+
"n": 15,
|
| 15341 |
"errors": 0,
|
| 15342 |
"error_rate": 0.0
|
| 15343 |
},
|
| 15344 |
"0.95": {
|
| 15345 |
+
"n": 14,
|
| 15346 |
"errors": 0,
|
| 15347 |
"error_rate": 0.0
|
| 15348 |
},
|
|
|
|
| 15355 |
"selective": {
|
| 15356 |
"0.5": {
|
| 15357 |
"coverage": 0.5,
|
| 15358 |
+
"accuracy": 1.0,
|
| 15359 |
+
"confidence_cutoff": 0.8030261229089575
|
| 15360 |
},
|
| 15361 |
"0.8": {
|
| 15362 |
"coverage": 0.8055555555555556,
|
| 15363 |
+
"accuracy": 0.7931034482758621,
|
| 15364 |
+
"confidence_cutoff": 0.5168578524128333
|
| 15365 |
}
|
| 15366 |
}
|
| 15367 |
}
|
|
|
|
| 15369 |
"heldout_tasks": {},
|
| 15370 |
"permutation": {
|
| 15371 |
"n": 36,
|
| 15372 |
+
"mean_max_delta": 0.07510801427468833,
|
| 15373 |
+
"flip_rate": 0.16666666666666666
|
| 15374 |
},
|
| 15375 |
"temperature": 1.0,
|
| 15376 |
"calibrated_clean": {
|
| 15377 |
"n": 1046,
|
| 15378 |
+
"nll": 0.9079076925806586,
|
| 15379 |
+
"acc": 0.6959847036328872,
|
| 15380 |
+
"ece": 0.11803573910263573,
|
| 15381 |
+
"brier": 0.41689550886050714,
|
| 15382 |
+
"mean_conf": 0.814020442735523,
|
| 15383 |
+
"confident_error_rate": 0.058317399617590825,
|
| 15384 |
+
"coverage_at_0_9": 0.5525812619502868,
|
| 15385 |
+
"accuracy_at_0_9": 0.8944636678200693,
|
| 15386 |
+
"coverage_at_5pct_error": 0.36424474187380496,
|
| 15387 |
+
"coverage_at_1pct_error": 0.14340344168260039,
|
| 15388 |
+
"aurc": 0.11088769274123257,
|
| 15389 |
+
"error_rate_at_0_9": 0.10553633217993079,
|
| 15390 |
+
"confidence_bias": 0.11803573910263576,
|
| 15391 |
"top_bins": {
|
| 15392 |
"0.9": {
|
| 15393 |
+
"n": 578,
|
| 15394 |
+
"errors": 61,
|
| 15395 |
+
"error_rate": 0.10553633217993079
|
| 15396 |
},
|
| 15397 |
"0.95": {
|
| 15398 |
+
"n": 491,
|
| 15399 |
+
"errors": 39,
|
| 15400 |
+
"error_rate": 0.07942973523421588
|
| 15401 |
},
|
| 15402 |
"0.99": {
|
| 15403 |
+
"n": 314,
|
| 15404 |
+
"errors": 11,
|
| 15405 |
+
"error_rate": 0.03503184713375796
|
| 15406 |
}
|
| 15407 |
},
|
| 15408 |
"selective": {
|
| 15409 |
"0.5": {
|
| 15410 |
"coverage": 0.5,
|
| 15411 |
+
"accuracy": 0.9101338432122371,
|
| 15412 |
+
"confidence_cutoff": 0.9301925551074537
|
| 15413 |
},
|
| 15414 |
"0.8": {
|
| 15415 |
"coverage": 0.8001912045889101,
|
| 15416 |
+
"accuracy": 0.7921146953405018,
|
| 15417 |
+
"confidence_cutoff": 0.5957019462220806
|
| 15418 |
}
|
| 15419 |
},
|
| 15420 |
"score_mae": 0.35059676214309027,
|
|
|
|
| 66859 |
}
|
| 66860 |
},
|
| 66861 |
"paired_comparisons": {
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 66862 |
"Sol minus kev-0.8b": {
|
| 66863 |
"old_core": {
|
| 66864 |
"delta": 0.1360863095238095,
|
|
|
|
| 67497 |
]
|
| 67498 |
}
|
| 67499 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 67500 |
"Lux minus Sol": {
|
| 67501 |
"old_core": {
|
| 67502 |
"delta": 0.09464285714285714,
|
|
|
|
| 68134 |
0.11044819701768911
|
| 68135 |
]
|
| 68136 |
}
|
| 68137 |
+
},
|
| 68138 |
+
"Nox minus kev-4b": {
|
| 68139 |
+
"old_core": {
|
| 68140 |
+
"delta": 0.11097470238095238,
|
| 68141 |
+
"exact_fraction": "2983/26880",
|
| 68142 |
+
"paired_ci95": [
|
| 68143 |
+
0.068154761904762,
|
| 68144 |
+
0.1539462425595238
|
| 68145 |
+
]
|
| 68146 |
+
},
|
| 68147 |
+
"v3_core": {
|
| 68148 |
+
"delta": 0.0325,
|
| 68149 |
+
"exact_fraction": "13/400",
|
| 68150 |
+
"paired_ci95": [
|
| 68151 |
+
0.0033333333333334103,
|
| 68152 |
+
0.06166666666666676
|
| 68153 |
+
]
|
| 68154 |
+
},
|
| 68155 |
+
"v4": {
|
| 68156 |
+
"delta": -0.028125,
|
| 68157 |
+
"exact_fraction": "-9/320",
|
| 68158 |
+
"paired_ci95": [
|
| 68159 |
+
-0.062443225931676956,
|
| 68160 |
+
0.00613059214504551
|
| 68161 |
+
]
|
| 68162 |
+
},
|
| 68163 |
+
"v5": {
|
| 68164 |
+
"delta": 0.016666666666666666,
|
| 68165 |
+
"exact_fraction": "1/60",
|
| 68166 |
+
"paired_ci95": [
|
| 68167 |
+
-0.014652826016486136,
|
| 68168 |
+
0.04855467675421749
|
| 68169 |
+
]
|
| 68170 |
+
},
|
| 68171 |
+
"transfer_v9_test": {
|
| 68172 |
+
"delta": -0.06500956022944551,
|
| 68173 |
+
"exact_fraction": "-34/523",
|
| 68174 |
+
"paired_ci95": [
|
| 68175 |
+
-0.09099616858237547,
|
| 68176 |
+
-0.03925188928140224
|
| 68177 |
+
]
|
| 68178 |
+
},
|
| 68179 |
+
"original_four_panel_mean": {
|
| 68180 |
+
"delta": 0.03300409226190476,
|
| 68181 |
+
"exact_fraction": "17743/537600",
|
| 68182 |
+
"paired_ci95": [
|
| 68183 |
+
0.0159171480237136,
|
| 68184 |
+
0.05071287650329332
|
| 68185 |
+
]
|
| 68186 |
+
},
|
| 68187 |
+
"weighted_mean": {
|
| 68188 |
+
"delta": 0.029947226679868887,
|
| 68189 |
+
"exact_fraction": "1403351/46860800",
|
| 68190 |
+
"paired_ci95": [
|
| 68191 |
+
0.013201872388473554,
|
| 68192 |
+
0.047038737697081195
|
| 68193 |
+
]
|
| 68194 |
+
}
|
| 68195 |
+
},
|
| 68196 |
+
"Nox minus Qwen3.5-4B": {
|
| 68197 |
+
"old_core": {
|
| 68198 |
+
"delta": 0.13113839285714285,
|
| 68199 |
+
"exact_fraction": "235/1792",
|
| 68200 |
+
"paired_ci95": [
|
| 68201 |
+
0.09270647321428567,
|
| 68202 |
+
0.16990420386904764
|
| 68203 |
+
]
|
| 68204 |
+
},
|
| 68205 |
+
"v3_core": {
|
| 68206 |
+
"delta": 0.08458333333333333,
|
| 68207 |
+
"exact_fraction": "203/2400",
|
| 68208 |
+
"paired_ci95": [
|
| 68209 |
+
0.05416666666666664,
|
| 68210 |
+
0.11541666666666661
|
| 68211 |
+
]
|
| 68212 |
+
},
|
| 68213 |
+
"v4": {
|
| 68214 |
+
"delta": -0.0890625,
|
| 68215 |
+
"exact_fraction": "-57/640",
|
| 68216 |
+
"paired_ci95": [
|
| 68217 |
+
-0.1342959349117049,
|
| 68218 |
+
-0.045753734276729595
|
| 68219 |
+
]
|
| 68220 |
+
},
|
| 68221 |
+
"v5": {
|
| 68222 |
+
"delta": 0.06458333333333334,
|
| 68223 |
+
"exact_fraction": "31/480",
|
| 68224 |
+
"paired_ci95": [
|
| 68225 |
+
0.028211805555555566,
|
| 68226 |
+
0.10191039756413821
|
| 68227 |
+
]
|
| 68228 |
+
},
|
| 68229 |
+
"transfer_v9_test": {
|
| 68230 |
+
"delta": 0.0076481835564053535,
|
| 68231 |
+
"exact_fraction": "4/523",
|
| 68232 |
+
"paired_ci95": [
|
| 68233 |
+
-0.022396190701526143,
|
| 68234 |
+
0.03838776086034498
|
| 68235 |
+
]
|
| 68236 |
+
},
|
| 68237 |
+
"original_four_panel_mean": {
|
| 68238 |
+
"delta": 0.04781063988095238,
|
| 68239 |
+
"exact_fraction": "25703/537600",
|
| 68240 |
+
"paired_ci95": [
|
| 68241 |
+
0.028947119327501308,
|
| 68242 |
+
0.06683696978735856
|
| 68243 |
+
]
|
| 68244 |
+
},
|
| 68245 |
+
"weighted_mean": {
|
| 68246 |
+
"delta": 0.05796270372393699,
|
| 68247 |
+
"exact_fraction": "1018567/17572800",
|
| 68248 |
+
"paired_ci95": [
|
| 68249 |
+
0.04083969253575949,
|
| 68250 |
+
0.07488554488575702
|
| 68251 |
+
]
|
| 68252 |
+
}
|
| 68253 |
+
},
|
| 68254 |
+
"Nox minus Decider": {
|
| 68255 |
+
"old_core": {
|
| 68256 |
+
"delta": 0.18988095238095237,
|
| 68257 |
+
"exact_fraction": "319/1680",
|
| 68258 |
+
"paired_ci95": [
|
| 68259 |
+
0.1443443080357141,
|
| 68260 |
+
0.23571428571428554
|
| 68261 |
+
]
|
| 68262 |
+
},
|
| 68263 |
+
"v3_core": {
|
| 68264 |
+
"delta": 0.052083333333333336,
|
| 68265 |
+
"exact_fraction": "5/96",
|
| 68266 |
+
"paired_ci95": [
|
| 68267 |
+
0.013333333333333308,
|
| 68268 |
+
0.08999999999999997
|
| 68269 |
+
]
|
| 68270 |
+
},
|
| 68271 |
+
"v4": {
|
| 68272 |
+
"delta": -0.1296875,
|
| 68273 |
+
"exact_fraction": "-83/640",
|
| 68274 |
+
"paired_ci95": [
|
| 68275 |
+
-0.17179976851851855,
|
| 68276 |
+
-0.08906249999999993
|
| 68277 |
+
]
|
| 68278 |
+
},
|
| 68279 |
+
"v5": {
|
| 68280 |
+
"delta": 0.01875,
|
| 68281 |
+
"exact_fraction": "3/160",
|
| 68282 |
+
"paired_ci95": [
|
| 68283 |
+
-0.012535333369549428,
|
| 68284 |
+
0.04969926075268807
|
| 68285 |
+
]
|
| 68286 |
+
},
|
| 68287 |
+
"transfer_v9_test": {
|
| 68288 |
+
"delta": 0.0028680688336520078,
|
| 68289 |
+
"exact_fraction": "3/1046",
|
| 68290 |
+
"paired_ci95": [
|
| 68291 |
+
-0.030160845668367766,
|
| 68292 |
+
0.03653934071222337
|
| 68293 |
+
]
|
| 68294 |
+
},
|
| 68295 |
+
"original_four_panel_mean": {
|
| 68296 |
+
"delta": 0.03275669642857143,
|
| 68297 |
+
"exact_fraction": "587/17920",
|
| 68298 |
+
"paired_ci95": [
|
| 68299 |
+
0.013155496386521998,
|
| 68300 |
+
0.05228677715137045
|
| 68301 |
+
]
|
| 68302 |
+
},
|
| 68303 |
+
"weighted_mean": {
|
| 68304 |
+
"delta": 0.05377470437266685,
|
| 68305 |
+
"exact_fraction": "7559777/140582400",
|
| 68306 |
+
"paired_ci95": [
|
| 68307 |
+
0.03467000712516513,
|
| 68308 |
+
0.07268914300878532
|
| 68309 |
+
]
|
| 68310 |
+
}
|
| 68311 |
}
|
| 68312 |
},
|
| 68313 |
"method": {
|
|
|
|
| 68335 |
"transfer_accuracy": "Upstream ordered probability argmax, clean knowable records, micro average. Refused/missing probabilities count wrong; unknown evidence is not accuracy.",
|
| 68336 |
"transfer_clusters": "Union IDs/group/parent/control/pair plus exact state SHA and ordered-request IDs, built over all variants then scored only clean knowable rows. Complete components jointly resampled; variable sampled denominator. No whole-family grouping.",
|
| 68337 |
"pairing": "Identical sampled source components across all models; five panels sampled independently. Original four marginal CIs preserved.",
|
| 68338 |
+
"accuracy_arithmetic": "Exact rational integer correct/requested; weights not fitted to outcomes.",
|
| 68339 |
"interval_role": "Descriptive paired95% intervals; no positive-CI release requirement invented.",
|
| 68340 |
"diagnostics": "Full per-task/proper-score/order/unknown-evidence results remain separate; no accuracy/latency/ECE unit mixing.",
|
| 68341 |
+
"exposure": "Observed regression tests; weighting user-approved after earlier results; not a blind prospective benchmark."
|
| 68342 |
},
|
| 68343 |
"prior_weighted_scores": {
|
| 68344 |
"Jev": {
|
|
|
|
| 68358 |
]
|
| 68359 |
},
|
| 68360 |
"Nox": {
|
| 68361 |
+
"point": 0.724150437750387,
|
| 68362 |
+
"exact_fraction": "203605613/281164800",
|
| 68363 |
+
"interval_status": "Not recomputed for alternate weighting; only the point is displayed."
|
|
|
|
|
|
|
|
|
|
| 68364 |
},
|
| 68365 |
"kev-9b": {
|
| 68366 |
"point": 0.7201455089684057,
|
|
|
|
| 68445 |
},
|
| 68446 |
"hosted_frontier": "All official calls returned jev-1.13.0; closed weight identity/input retention cannot be independently inspected.",
|
| 68447 |
"size_note": "Qwen3.5 family size labels name their official parent checkpoint. Lux deployed text+decision weights total7,940,895,744 parameters; no vision or generativeLMhead in its serving bundle.",
|
| 68448 |
+
"presentation_note": "Latest released result per model. Nox uses its qualified Choice null-description rendering; other displayed model results are unchanged. No historical model rows.",
|
| 68449 |
+
"update_kind": "Nox SystemOne Choice semantics: null descriptions use their key text; model weights, tokenizer and temperature unchanged."
|
| 68450 |
}
|
metrics/evaluation-provenance.json
CHANGED
|
@@ -1,145 +1,148 @@
|
|
| 1 |
{
|
| 2 |
-
"
|
| 3 |
-
"
|
|
|
|
| 4 |
"qualified_runtime": {
|
| 5 |
"python": "3.12.13",
|
| 6 |
"numpy": "2.3.5"
|
| 7 |
},
|
| 8 |
-
"source_sha256": "
|
| 9 |
"model_manifest": {
|
| 10 |
-
"
|
| 11 |
"panels": {
|
| 12 |
"old_core": {
|
| 13 |
-
"path": "
|
| 14 |
-
"sha256": "
|
|
|
|
| 15 |
},
|
| 16 |
"v3_core": {
|
| 17 |
-
"path": "
|
| 18 |
-
"sha256": "
|
|
|
|
| 19 |
},
|
| 20 |
"v4": {
|
| 21 |
-
"path": "
|
| 22 |
-
"sha256": "
|
|
|
|
| 23 |
},
|
| 24 |
"v5": {
|
| 25 |
-
"path": "
|
| 26 |
-
"sha256": "
|
|
|
|
| 27 |
}
|
| 28 |
},
|
| 29 |
"transfer_rows": {
|
| 30 |
-
"path": "analysis/
|
| 31 |
-
"sha256": "
|
| 32 |
},
|
| 33 |
"evidence": [
|
| 34 |
{
|
| 35 |
-
"path": "
|
| 36 |
-
"sha256": "
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
},
|
| 38 |
{
|
| 39 |
-
"path": "analysis/
|
| 40 |
-
"sha256": "
|
| 41 |
}
|
| 42 |
]
|
| 43 |
},
|
| 44 |
-
"
|
| 45 |
"panels": {
|
| 46 |
"old_core": {
|
| 47 |
-
"path": "results/
|
| 48 |
-
"sha256": "
|
| 49 |
},
|
| 50 |
"v3_core": {
|
| 51 |
-
"path": "results/
|
| 52 |
-
"sha256": "
|
| 53 |
},
|
| 54 |
"v4": {
|
| 55 |
-
"path": "results/
|
| 56 |
-
"sha256": "
|
| 57 |
},
|
| 58 |
"v5": {
|
| 59 |
-
"path": "results/
|
| 60 |
-
"sha256": "
|
| 61 |
}
|
| 62 |
},
|
| 63 |
"transfer_rows": {
|
| 64 |
-
"path": "analysis/
|
| 65 |
-
"sha256": "
|
| 66 |
},
|
| 67 |
"evidence": [
|
| 68 |
{
|
| 69 |
-
"path": "results/
|
| 70 |
-
"sha256": "
|
| 71 |
},
|
| 72 |
{
|
| 73 |
-
"path": "
|
| 74 |
-
"sha256": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 75 |
}
|
| 76 |
]
|
| 77 |
},
|
| 78 |
-
"
|
| 79 |
"panels": {
|
| 80 |
"old_core": {
|
| 81 |
-
"path": "
|
| 82 |
-
"sha256": "
|
| 83 |
},
|
| 84 |
"v3_core": {
|
| 85 |
-
"path": "
|
| 86 |
-
"sha256": "
|
| 87 |
},
|
| 88 |
"v4": {
|
| 89 |
-
"path": "
|
| 90 |
-
"sha256": "
|
| 91 |
},
|
| 92 |
"v5": {
|
| 93 |
-
"path": "
|
| 94 |
-
"sha256": "
|
| 95 |
}
|
| 96 |
},
|
| 97 |
"transfer_rows": {
|
| 98 |
-
"path": "analysis/
|
| 99 |
-
"sha256": "
|
| 100 |
},
|
| 101 |
"evidence": [
|
| 102 |
{
|
| 103 |
-
"path": "
|
| 104 |
-
"sha256": "
|
| 105 |
},
|
| 106 |
{
|
| 107 |
-
"path": "
|
| 108 |
-
"sha256": "
|
| 109 |
-
}
|
| 110 |
-
]
|
| 111 |
-
},
|
| 112 |
-
"kev-4b": {
|
| 113 |
-
"panels": {
|
| 114 |
-
"old_core": {
|
| 115 |
-
"path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-old_core-normalized.jsonl",
|
| 116 |
-
"sha256": "41dfbb7c9c4381effc219f9123b51bd4d688b3e746cec974ef930730b9dfd6b1"
|
| 117 |
},
|
| 118 |
-
|
| 119 |
-
"path": "
|
| 120 |
-
"sha256": "
|
| 121 |
},
|
| 122 |
-
|
| 123 |
-
"path": "
|
| 124 |
-
"sha256": "
|
| 125 |
},
|
| 126 |
-
"v5": {
|
| 127 |
-
"path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v5-normalized.jsonl",
|
| 128 |
-
"sha256": "d33bbb60134314dcb550060176ed0a3ae91ec0c6b033cceb40e3255b33666898"
|
| 129 |
-
}
|
| 130 |
-
},
|
| 131 |
-
"transfer_rows": {
|
| 132 |
-
"path": "analysis/kev-reciprocal-v1/reports/kev-4b/v9-transfer-v9-test-published-temperature-rows.json",
|
| 133 |
-
"sha256": "c2cb16a936be4c10ca6b1870c69426ddd6b51629c401f71ae1da652b3bd2ac7e"
|
| 134 |
-
},
|
| 135 |
-
"evidence": [
|
| 136 |
{
|
| 137 |
-
"path": "
|
| 138 |
-
"sha256": "
|
| 139 |
},
|
| 140 |
{
|
| 141 |
-
"path": "analysis/
|
| 142 |
-
"sha256": "
|
| 143 |
}
|
| 144 |
]
|
| 145 |
},
|
|
@@ -177,129 +180,37 @@
|
|
| 177 |
}
|
| 178 |
]
|
| 179 |
},
|
| 180 |
-
"
|
| 181 |
-
"panels": {
|
| 182 |
-
"old_core": {
|
| 183 |
-
"path": "eval/heldout/official-predictions.jsonl",
|
| 184 |
-
"sha256": "069e48e1fdf1b5f3d95cc0097e4179ac3faab97bb357b0ce897cfb463b87cbaf",
|
| 185 |
-
"bytes": 185263
|
| 186 |
-
},
|
| 187 |
-
"v3_core": {
|
| 188 |
-
"path": "eval/heldout/v3/official/normalized/core.jsonl",
|
| 189 |
-
"sha256": "54707354bfe5d96b5de56e6dc1b38cba0a610eea516b28fbff013b94f422b397",
|
| 190 |
-
"bytes": 368716
|
| 191 |
-
},
|
| 192 |
-
"v4": {
|
| 193 |
-
"path": "eval/heldout/v4/official/normalized/core.jsonl",
|
| 194 |
-
"sha256": "d6583f903e974d70747fae2944014828954dc4d7344ceddcbb2cd1b06fbfc799",
|
| 195 |
-
"bytes": 183671
|
| 196 |
-
},
|
| 197 |
-
"v5": {
|
| 198 |
-
"path": "eval/heldout/v5/official/normalized/core.jsonl",
|
| 199 |
-
"sha256": "f883164c4c83db8e42ebe03f7e165afbd2f81f2b541bcf8008a2cc06a60e1ee7",
|
| 200 |
-
"bytes": 200316
|
| 201 |
-
}
|
| 202 |
-
},
|
| 203 |
-
"transfer_rows": {
|
| 204 |
-
"path": "analysis/decision-benchmark-v3-official/scored/ROWS.json",
|
| 205 |
-
"sha256": "be29a7f3804ad84993ffb06f05ada7e2d8ed25317504c4fc17fe5dc306fd788c"
|
| 206 |
-
},
|
| 207 |
-
"evidence": [
|
| 208 |
-
{
|
| 209 |
-
"path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
|
| 210 |
-
"sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
|
| 211 |
-
},
|
| 212 |
-
{
|
| 213 |
-
"path": "analysis/decision-benchmark-v3-official/scored/REPORT.json",
|
| 214 |
-
"sha256": "2fdbaf7db1e4b877762e33a87b23a236a065eaa3078b14c2d7d579ad44f9ea24"
|
| 215 |
-
},
|
| 216 |
-
{
|
| 217 |
-
"path": "analysis/decision-benchmark-v3-official/INDEPENDENT-PROJECTION-REVIEW.json",
|
| 218 |
-
"sha256": "b9506539e2c4a863aef16f7dab0517c175e85863416bc622c60df05489a9eb52"
|
| 219 |
-
}
|
| 220 |
-
]
|
| 221 |
-
},
|
| 222 |
-
"Laya-base": {
|
| 223 |
"panels": {
|
| 224 |
-
"v4": {
|
| 225 |
-
"path": "results/public-roster-baselines-v1/laya-english/worker/v4/normalized.jsonl",
|
| 226 |
-
"sha256": "31da1d4140da69d2677172707eb301843cf8fdfb43200cba93be910321f922ba"
|
| 227 |
-
},
|
| 228 |
-
"v5": {
|
| 229 |
-
"path": "results/public-roster-baselines-v1/laya-english/worker/v5/normalized.jsonl",
|
| 230 |
-
"sha256": "159614ceede814b578e1e025836a595c8c71a9bec4896caf6454d7bf54c4c44d"
|
| 231 |
-
},
|
| 232 |
"old_core": {
|
| 233 |
-
"path": "
|
| 234 |
-
"sha256": "
|
| 235 |
},
|
| 236 |
"v3_core": {
|
| 237 |
-
"path": "
|
| 238 |
-
"sha256": "
|
| 239 |
-
}
|
| 240 |
-
},
|
| 241 |
-
"transfer_rows": {
|
| 242 |
-
"path": "analysis/decision-benchmark-v3-baselines/laya-english/ROWS.json",
|
| 243 |
-
"sha256": "19138909259400facb5d3d5402d0d736f870695fe3ee46019f2e438088c57275"
|
| 244 |
-
},
|
| 245 |
-
"evidence": [
|
| 246 |
-
{
|
| 247 |
-
"path": "results/public-roster-baselines-v1/laya-english/worker/COMPLETE.json",
|
| 248 |
-
"sha256": "cef0da797ae14cebaf1e4b9edf9b16c3cd9a412a30435d7d2517b8ba6dbe8b98"
|
| 249 |
-
},
|
| 250 |
-
{
|
| 251 |
-
"path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
|
| 252 |
-
"sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
|
| 253 |
-
},
|
| 254 |
-
{
|
| 255 |
-
"path": "analysis/decision-benchmark-v3-baselines/laya-english/REPORT.json",
|
| 256 |
-
"sha256": "cb5935bb399dcf23ffe24f855452b2e074bcbb3d023cfc7dd439342e54f02c6b"
|
| 257 |
},
|
| 258 |
-
{
|
| 259 |
-
"path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
|
| 260 |
-
"sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
|
| 261 |
-
}
|
| 262 |
-
]
|
| 263 |
-
},
|
| 264 |
-
"Laya-multilingual": {
|
| 265 |
-
"panels": {
|
| 266 |
"v4": {
|
| 267 |
-
"path": "
|
| 268 |
-
"sha256": "
|
| 269 |
},
|
| 270 |
"v5": {
|
| 271 |
-
"path": "
|
| 272 |
-
"sha256": "
|
| 273 |
-
},
|
| 274 |
-
"old_core": {
|
| 275 |
-
"path": "eval/heldout/primary-v2/laya-multilingual/core/normalized.jsonl",
|
| 276 |
-
"sha256": "9ea39841c4d5901cd9ab860c980d2f3b55358be289c7e01b9b47982927bf849d"
|
| 277 |
-
},
|
| 278 |
-
"v3_core": {
|
| 279 |
-
"path": "eval/heldout/v3/baselines/laya-multilingual/core/normalized.jsonl",
|
| 280 |
-
"sha256": "a0b7ce9daf4be2c7831bcab6b5e96c8dcb6bef133d39475f72241ab18faccdd2"
|
| 281 |
}
|
| 282 |
},
|
| 283 |
"transfer_rows": {
|
| 284 |
-
"path": "analysis/
|
| 285 |
-
"sha256": "
|
| 286 |
},
|
| 287 |
"evidence": [
|
| 288 |
{
|
| 289 |
-
"path": "
|
| 290 |
-
"sha256": "
|
| 291 |
-
},
|
| 292 |
-
{
|
| 293 |
-
"path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
|
| 294 |
-
"sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
|
| 295 |
-
},
|
| 296 |
-
{
|
| 297 |
-
"path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/REPORT.json",
|
| 298 |
-
"sha256": "d20d3d58f76ca2758545f83d358abd747f86d581c897d1fb521371d1000d6154"
|
| 299 |
},
|
| 300 |
{
|
| 301 |
-
"path": "analysis/
|
| 302 |
-
"sha256": "
|
| 303 |
}
|
| 304 |
]
|
| 305 |
},
|
|
@@ -361,99 +272,171 @@
|
|
| 361 |
}
|
| 362 |
]
|
| 363 |
},
|
| 364 |
-
"
|
| 365 |
"panels": {
|
| 366 |
"old_core": {
|
| 367 |
-
"path": "
|
| 368 |
-
"sha256": "
|
| 369 |
},
|
| 370 |
"v3_core": {
|
| 371 |
-
"path": "
|
| 372 |
-
"sha256": "
|
| 373 |
},
|
| 374 |
"v4": {
|
| 375 |
-
"path": "
|
| 376 |
-
"sha256": "
|
| 377 |
},
|
| 378 |
"v5": {
|
| 379 |
-
"path": "
|
| 380 |
-
"sha256": "
|
| 381 |
}
|
| 382 |
},
|
| 383 |
"transfer_rows": {
|
| 384 |
-
"path": "analysis/decision-benchmark-v3-baselines/
|
| 385 |
-
"sha256": "
|
| 386 |
},
|
| 387 |
"evidence": [
|
| 388 |
{
|
| 389 |
-
"path": "results/
|
| 390 |
-
"sha256": "
|
| 391 |
},
|
| 392 |
{
|
| 393 |
-
"path": "results/
|
| 394 |
-
"sha256": "
|
| 395 |
},
|
| 396 |
{
|
| 397 |
-
"path": "
|
| 398 |
-
"sha256": "
|
| 399 |
},
|
| 400 |
{
|
| 401 |
-
"path": "analysis/decision-benchmark-v3-baselines/
|
| 402 |
-
"sha256": "
|
| 403 |
},
|
| 404 |
{
|
| 405 |
-
"path": "
|
| 406 |
-
"sha256": "
|
|
|
|
|
|
|
|
|
|
|
|
|
| 407 |
}
|
| 408 |
]
|
| 409 |
},
|
| 410 |
-
"
|
| 411 |
"panels": {
|
| 412 |
"old_core": {
|
| 413 |
-
"path": "eval/heldout/
|
| 414 |
-
"sha256": "
|
| 415 |
},
|
| 416 |
"v3_core": {
|
| 417 |
-
"path": "eval/heldout/v3/baselines/
|
| 418 |
-
"sha256": "
|
| 419 |
},
|
| 420 |
"v4": {
|
| 421 |
-
"path": "eval/heldout/v4/baselines/
|
| 422 |
-
"sha256": "
|
| 423 |
},
|
| 424 |
"v5": {
|
| 425 |
-
"path": "eval/heldout/v5/open-baselines-v1/
|
| 426 |
-
"sha256": "
|
| 427 |
}
|
| 428 |
},
|
| 429 |
"transfer_rows": {
|
| 430 |
-
"path": "analysis/decision-benchmark-v3-baselines/
|
| 431 |
-
"sha256": "
|
| 432 |
},
|
| 433 |
"evidence": [
|
| 434 |
{
|
| 435 |
-
"path": "results/
|
| 436 |
-
"sha256": "
|
| 437 |
},
|
| 438 |
{
|
| 439 |
-
"path": "results/
|
| 440 |
-
"sha256": "
|
| 441 |
},
|
| 442 |
{
|
| 443 |
-
"path": "results/
|
| 444 |
-
"sha256": "
|
| 445 |
},
|
| 446 |
{
|
| 447 |
-
"path": "analysis/decision-benchmark-v3-baselines/
|
| 448 |
-
"sha256": "
|
| 449 |
},
|
| 450 |
{
|
| 451 |
"path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
|
| 452 |
"sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
|
| 453 |
},
|
| 454 |
{
|
| 455 |
-
"path": "analysis/decision-benchmark-v3-baselines/
|
| 456 |
-
"sha256": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 457 |
}
|
| 458 |
]
|
| 459 |
},
|
|
@@ -507,60 +490,93 @@
|
|
| 507 |
}
|
| 508 |
]
|
| 509 |
},
|
| 510 |
-
"
|
| 511 |
"panels": {
|
| 512 |
-
"old_core": {
|
| 513 |
-
"path": "eval/heldout/primary-v2/base-4b/core/normalized.jsonl",
|
| 514 |
-
"sha256": "d20ad845066ff81caaf60ebee44df941872bdc642da78b5f3b3da4a33dd9a08d"
|
| 515 |
-
},
|
| 516 |
-
"v3_core": {
|
| 517 |
-
"path": "eval/heldout/v3/baselines/base-4b/core/normalized.jsonl",
|
| 518 |
-
"sha256": "7541893d2b2ec0ec2ffc513f50c568eaef860b65bcef9a1747996c116fff9661"
|
| 519 |
-
},
|
| 520 |
"v4": {
|
| 521 |
-
"path": "
|
| 522 |
-
"sha256": "
|
| 523 |
},
|
| 524 |
"v5": {
|
| 525 |
-
"path": "
|
| 526 |
-
"sha256": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 527 |
}
|
| 528 |
},
|
| 529 |
"transfer_rows": {
|
| 530 |
-
"path": "analysis/decision-benchmark-v3-baselines/
|
| 531 |
-
"sha256": "
|
| 532 |
},
|
| 533 |
"evidence": [
|
| 534 |
{
|
| 535 |
-
"path": "results/
|
| 536 |
-
"sha256": "
|
| 537 |
},
|
| 538 |
{
|
| 539 |
-
"path": "
|
| 540 |
-
"sha256": "
|
| 541 |
},
|
| 542 |
{
|
| 543 |
-
"path": "
|
| 544 |
-
"sha256": "
|
| 545 |
},
|
| 546 |
{
|
| 547 |
-
"path": "analysis/
|
| 548 |
-
"sha256": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 549 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 550 |
{
|
| 551 |
-
"path": "
|
| 552 |
-
"sha256": "
|
| 553 |
},
|
| 554 |
{
|
| 555 |
-
"path": "analysis/
|
| 556 |
-
"sha256": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 557 |
}
|
| 558 |
]
|
| 559 |
}
|
| 560 |
},
|
| 561 |
"task_rows": 54,
|
| 562 |
"rank_public_models": [
|
| 563 |
-
"Jev",
|
| 564 |
"Lux",
|
| 565 |
"Nox",
|
| 566 |
"kev-9b",
|
|
@@ -572,7 +588,8 @@
|
|
| 572 |
"kev-0.8b",
|
| 573 |
"Qwen3.5-2B",
|
| 574 |
"Laya-base",
|
| 575 |
-
"Laya-multilingual"
|
|
|
|
| 576 |
],
|
| 577 |
"internal_extra_models_not_public_rank": [
|
| 578 |
"llm2jev-2b",
|
|
@@ -655,7 +672,5 @@
|
|
| 655 |
}
|
| 656 |
},
|
| 657 |
"source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
|
| 658 |
-
}
|
| 659 |
-
"frozen_protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
|
| 660 |
-
"presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
|
| 661 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"frozen_protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
|
| 3 |
+
"statistics_sha256": "b319a993ac569ea1fa8da027ccb2195868557e4efe232a5d06091db2884086be",
|
| 4 |
+
"manifest_sha256": "cf455486f8eba2b7158bbb4b5bcdcd8a4bd4daee4f05184b1df61b5f70ef2b04",
|
| 5 |
"qualified_runtime": {
|
| 6 |
"python": "3.12.13",
|
| 7 |
"numpy": "2.3.5"
|
| 8 |
},
|
| 9 |
+
"source_sha256": "f2f72fe161944599b1114df29ff095786bf661cca59afd9d784b7ca6d1430c99",
|
| 10 |
"model_manifest": {
|
| 11 |
+
"Jev": {
|
| 12 |
"panels": {
|
| 13 |
"old_core": {
|
| 14 |
+
"path": "eval/heldout/official-predictions.jsonl",
|
| 15 |
+
"sha256": "069e48e1fdf1b5f3d95cc0097e4179ac3faab97bb357b0ce897cfb463b87cbaf",
|
| 16 |
+
"bytes": 185263
|
| 17 |
},
|
| 18 |
"v3_core": {
|
| 19 |
+
"path": "eval/heldout/v3/official/normalized/core.jsonl",
|
| 20 |
+
"sha256": "54707354bfe5d96b5de56e6dc1b38cba0a610eea516b28fbff013b94f422b397",
|
| 21 |
+
"bytes": 368716
|
| 22 |
},
|
| 23 |
"v4": {
|
| 24 |
+
"path": "eval/heldout/v4/official/normalized/core.jsonl",
|
| 25 |
+
"sha256": "d6583f903e974d70747fae2944014828954dc4d7344ceddcbb2cd1b06fbfc799",
|
| 26 |
+
"bytes": 183671
|
| 27 |
},
|
| 28 |
"v5": {
|
| 29 |
+
"path": "eval/heldout/v5/official/normalized/core.jsonl",
|
| 30 |
+
"sha256": "f883164c4c83db8e42ebe03f7e165afbd2f81f2b541bcf8008a2cc06a60e1ee7",
|
| 31 |
+
"bytes": 200316
|
| 32 |
}
|
| 33 |
},
|
| 34 |
"transfer_rows": {
|
| 35 |
+
"path": "analysis/decision-benchmark-v3-official/scored/ROWS.json",
|
| 36 |
+
"sha256": "be29a7f3804ad84993ffb06f05ada7e2d8ed25317504c4fc17fe5dc306fd788c"
|
| 37 |
},
|
| 38 |
"evidence": [
|
| 39 |
{
|
| 40 |
+
"path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
|
| 41 |
+
"sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
|
| 42 |
+
},
|
| 43 |
+
{
|
| 44 |
+
"path": "analysis/decision-benchmark-v3-official/scored/REPORT.json",
|
| 45 |
+
"sha256": "2fdbaf7db1e4b877762e33a87b23a236a065eaa3078b14c2d7d579ad44f9ea24"
|
| 46 |
},
|
| 47 |
{
|
| 48 |
+
"path": "analysis/decision-benchmark-v3-official/INDEPENDENT-PROJECTION-REVIEW.json",
|
| 49 |
+
"sha256": "b9506539e2c4a863aef16f7dab0517c175e85863416bc622c60df05489a9eb52"
|
| 50 |
}
|
| 51 |
]
|
| 52 |
},
|
| 53 |
+
"Lux": {
|
| 54 |
"panels": {
|
| 55 |
"old_core": {
|
| 56 |
+
"path": "results/lux-new-host-v1/quality/lux/old_core/normalized.jsonl",
|
| 57 |
+
"sha256": "a659a7c7fef2871bbf806e842ea8ca81f645dcf610f0f2f6622d8d32e76a0836"
|
| 58 |
},
|
| 59 |
"v3_core": {
|
| 60 |
+
"path": "results/lux-new-host-v1/quality/lux/v3_core/normalized.jsonl",
|
| 61 |
+
"sha256": "da8903a07e36a9687f6ee1d0fd8dbc684e61c49ef5fea30c346e46bff5a90e94"
|
| 62 |
},
|
| 63 |
"v4": {
|
| 64 |
+
"path": "results/lux-new-host-v1/quality/lux/v4/normalized.jsonl",
|
| 65 |
+
"sha256": "26cd7482693ec1ffa1a93a156abfc88378e4872e320f3b9b775763635086d488"
|
| 66 |
},
|
| 67 |
"v5": {
|
| 68 |
+
"path": "results/lux-new-host-v1/quality/lux/v5/normalized.jsonl",
|
| 69 |
+
"sha256": "759c773e629a16317271de850780a8734533c795046122f850914c4482d1e72b"
|
| 70 |
}
|
| 71 |
},
|
| 72 |
"transfer_rows": {
|
| 73 |
+
"path": "analysis/decision-benchmark-v3-baselines/Lux/ROWS.json",
|
| 74 |
+
"sha256": "bafb55fc86762061e537382360eda14a8a8f8c5e4e111c2fda1f82cd6c06793f"
|
| 75 |
},
|
| 76 |
"evidence": [
|
| 77 |
{
|
| 78 |
+
"path": "results/lux-new-host-v1/quality/lux/COMPLETE.json",
|
| 79 |
+
"sha256": "9600ff792cbda13af5505c804cf63af94d2ee895f2912f0645a7d493ae1b7a9c"
|
| 80 |
},
|
| 81 |
{
|
| 82 |
+
"path": "results/lux-new-host-v1/quality/lux/metadata.json",
|
| 83 |
+
"sha256": "0d4509fae7eb3630118eaebc24e75a7b02f717d9a0ad74811fa27358c33c9cc5"
|
| 84 |
+
},
|
| 85 |
+
{
|
| 86 |
+
"path": "analysis/decision-benchmark-v3-baselines/Lux-projected/COMPLETE.json",
|
| 87 |
+
"sha256": "f58d02df78f297d4feebe39facfac072951111d8d6a19e2c1069f1b90250c622"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"path": "analysis/decision-benchmark-v3-baselines/Lux/REPORT.json",
|
| 91 |
+
"sha256": "f31703c342efe043a1e1534da189808bb7e0de0641fbaf3f91faa4a1f928a299"
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"path": "analysis/decoder4b/lux9b-training-v1/CHECKPOINT-HELDOUT-COLLECTED.json",
|
| 95 |
+
"sha256": "d1c42515ab225b27ea011a6cf2262fe12dd7001e3b4d2cb64a23e7642c7a0d91"
|
| 96 |
}
|
| 97 |
]
|
| 98 |
},
|
| 99 |
+
"Nox": {
|
| 100 |
"panels": {
|
| 101 |
"old_core": {
|
| 102 |
+
"path": "results/nox-null-description-v1/full-regression/old_core/normalized.jsonl",
|
| 103 |
+
"sha256": "dbd18273dbbb9d09ba37d19dfedb0fcd6b64841d528632f04d39fd7955799967"
|
| 104 |
},
|
| 105 |
"v3_core": {
|
| 106 |
+
"path": "results/nox-null-description-v1/full-regression/v3_core/normalized.jsonl",
|
| 107 |
+
"sha256": "5286d1c3a78649e3284a018b1518f94cf32d8f0ad43a49c6a14c993524a3aff8"
|
| 108 |
},
|
| 109 |
"v4": {
|
| 110 |
+
"path": "results/nox-null-description-v1/full-regression/v4/normalized.jsonl",
|
| 111 |
+
"sha256": "077b752f4866bd0fc83db9b31c6227fd5015ce3585d2138c1a7a803c433d472b"
|
| 112 |
},
|
| 113 |
"v5": {
|
| 114 |
+
"path": "results/nox-null-description-v1/full-regression/v5/normalized.jsonl",
|
| 115 |
+
"sha256": "d57f07df59ff9b0e6ace0658e7917746e2b8d3bf518309a3858045cd175d79df"
|
| 116 |
}
|
| 117 |
},
|
| 118 |
"transfer_rows": {
|
| 119 |
+
"path": "analysis/nox-null-description-candidate-v4/transfer/ROWS.json",
|
| 120 |
+
"sha256": "7cee9e036ee53a3c41918cfafc73ad0d68f16b266abcae43077d1d69247afdfe"
|
| 121 |
},
|
| 122 |
"evidence": [
|
| 123 |
{
|
| 124 |
+
"path": "results/nox-null-description-v1/COMPLETE.json",
|
| 125 |
+
"sha256": "b813aa5db8c43b728b94abbaef20cefc8d3d84b51f7873ca46b3afb59aab6fe4"
|
| 126 |
},
|
| 127 |
{
|
| 128 |
+
"path": "results/nox-null-description-v1/full-regression/COMPLETE.json",
|
| 129 |
+
"sha256": "ad80e5bfb2558ee3b8fb1fc80617fcefbc1fbb8e58b8155e3cf5dfde38528754"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 130 |
},
|
| 131 |
+
{
|
| 132 |
+
"path": "results/nox-null-description-v1/transfer-v9/COMPLETE.json",
|
| 133 |
+
"sha256": "93a8767f7245b93764e527e616f236b5fa46b24a087a7e8ff66d35eee9f9fd0f"
|
| 134 |
},
|
| 135 |
+
{
|
| 136 |
+
"path": "results/nox-null-description-v1/full-regression-ACTUAL-EXIT.json",
|
| 137 |
+
"sha256": "dca6e05615c9c82a2d506438fb4e7c9abfe03053661515ffc9a4c5b176fd3ed8"
|
| 138 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 139 |
{
|
| 140 |
+
"path": "results/nox-null-description-v1/transfer-v9-ACTUAL-EXIT.json",
|
| 141 |
+
"sha256": "09c4c05549253c386f72b52e8ad4fbbc47d05af59715caebf9d01a6e8734d0af"
|
| 142 |
},
|
| 143 |
{
|
| 144 |
+
"path": "analysis/nox-null-description-candidate-v4/transfer/REPORT.json",
|
| 145 |
+
"sha256": "f02e447787c0c5afcb8acbd06408407719ebfc191e8501e8a887cd96c150f81b"
|
| 146 |
}
|
| 147 |
]
|
| 148 |
},
|
|
|
|
| 180 |
}
|
| 181 |
]
|
| 182 |
},
|
| 183 |
+
"kev-4b": {
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 184 |
"panels": {
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 185 |
"old_core": {
|
| 186 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-old_core-normalized.jsonl",
|
| 187 |
+
"sha256": "41dfbb7c9c4381effc219f9123b51bd4d688b3e746cec974ef930730b9dfd6b1"
|
| 188 |
},
|
| 189 |
"v3_core": {
|
| 190 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v3_core-normalized.jsonl",
|
| 191 |
+
"sha256": "2ec54da92c4e72412e89b3c07013b54729f5899ba8542eeadf962b04a1f3a4b8"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 192 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 193 |
"v4": {
|
| 194 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v4-normalized.jsonl",
|
| 195 |
+
"sha256": "88c576c3d62ec11686dd5ebf360971a9087779e45a346897e2badf874a754a12"
|
| 196 |
},
|
| 197 |
"v5": {
|
| 198 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v5-normalized.jsonl",
|
| 199 |
+
"sha256": "d33bbb60134314dcb550060176ed0a3ae91ec0c6b033cceb40e3255b33666898"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 200 |
}
|
| 201 |
},
|
| 202 |
"transfer_rows": {
|
| 203 |
+
"path": "analysis/kev-reciprocal-v1/reports/kev-4b/v9-transfer-v9-test-published-temperature-rows.json",
|
| 204 |
+
"sha256": "c2cb16a936be4c10ca6b1870c69426ddd6b51629c401f71ae1da652b3bd2ac7e"
|
| 205 |
},
|
| 206 |
"evidence": [
|
| 207 |
{
|
| 208 |
+
"path": "analysis/kev-reciprocal-v1/reports/kev-4b/REPORT.json",
|
| 209 |
+
"sha256": "366fb0d27a4d732e5e270c264521d9a93ccabc4f0ddf5b7d3e9f1ecc5f460406"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 210 |
},
|
| 211 |
{
|
| 212 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
|
| 213 |
+
"sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
|
| 214 |
}
|
| 215 |
]
|
| 216 |
},
|
|
|
|
| 272 |
}
|
| 273 |
]
|
| 274 |
},
|
| 275 |
+
"Decider": {
|
| 276 |
"panels": {
|
| 277 |
"old_core": {
|
| 278 |
+
"path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/decider/quality/normalized.jsonl",
|
| 279 |
+
"sha256": "e1b14f0f9ec521bffbde5945684bc2410a413b6b6418d65ac4fc702a7b5e7937"
|
| 280 |
},
|
| 281 |
"v3_core": {
|
| 282 |
+
"path": "eval/heldout/v3/baselines/decider/core/normalized.jsonl",
|
| 283 |
+
"sha256": "12a1f554decf1aff608743e7b4a44681289ea084ae25f383e519b1f34917d46f"
|
| 284 |
},
|
| 285 |
"v4": {
|
| 286 |
+
"path": "eval/heldout/v4/baselines/decider/core/normalized.jsonl",
|
| 287 |
+
"sha256": "626ef366ae471ec10dfb89ef2ff2f7b29a6c879976d3ff14fcbb2bafd7d9b045"
|
| 288 |
},
|
| 289 |
"v5": {
|
| 290 |
+
"path": "eval/heldout/v5/open-baselines-v1/decider/core/normalized.jsonl",
|
| 291 |
+
"sha256": "ae16b17ac8e4208940c0c3e9e04f25e94686b5d16bf46f64f53669baf41aeb71"
|
| 292 |
}
|
| 293 |
},
|
| 294 |
"transfer_rows": {
|
| 295 |
+
"path": "analysis/decision-benchmark-v3-baselines/decider/ROWS.json",
|
| 296 |
+
"sha256": "6da5b5692a8a0cea3dc8f6d02d911fafd9bbf62e286a9be8db27c89f77a897e4"
|
| 297 |
},
|
| 298 |
"evidence": [
|
| 299 |
{
|
| 300 |
+
"path": "results/accelerated-transfer-baselines-v1/decider/worker/COMPLETE.json",
|
| 301 |
+
"sha256": "e1ee0df8df59ec3eb2b49ba947a8508994a24bce226b7f2cccbe25f6835b99af"
|
| 302 |
},
|
| 303 |
{
|
| 304 |
+
"path": "results/accelerated-transfer-baselines-v1/decider/worker/METADATA.json",
|
| 305 |
+
"sha256": "f00061c1b0bb642bf1353aae867d19df988c01b0428b58206b31249521bcbff8"
|
| 306 |
},
|
| 307 |
{
|
| 308 |
+
"path": "results/accelerated-transfer-baselines-v1/decider/worker/predictions.jsonl",
|
| 309 |
+
"sha256": "f69e87fa65b09289e6861c1d789270aea5a92d23ac47817342647ed08f406df7"
|
| 310 |
},
|
| 311 |
{
|
| 312 |
+
"path": "analysis/decision-benchmark-v3-baselines/decider/REPORT.json",
|
| 313 |
+
"sha256": "5aae029f34524dfebc8e79128da8c39229c5747461dfce07ca59956da4416951"
|
| 314 |
},
|
| 315 |
{
|
| 316 |
+
"path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
|
| 317 |
+
"sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
|
| 318 |
+
},
|
| 319 |
+
{
|
| 320 |
+
"path": "analysis/decision-benchmark-v3-baselines/decider/COMPOSABLE-MANIFEST.json",
|
| 321 |
+
"sha256": "a49848e86feafaa4f36a33287910035e7605e344791aae139305b78cb304a77f"
|
| 322 |
}
|
| 323 |
]
|
| 324 |
},
|
| 325 |
+
"Qwen3.5-4B": {
|
| 326 |
"panels": {
|
| 327 |
"old_core": {
|
| 328 |
+
"path": "eval/heldout/primary-v2/base-4b/core/normalized.jsonl",
|
| 329 |
+
"sha256": "d20ad845066ff81caaf60ebee44df941872bdc642da78b5f3b3da4a33dd9a08d"
|
| 330 |
},
|
| 331 |
"v3_core": {
|
| 332 |
+
"path": "eval/heldout/v3/baselines/base-4b/core/normalized.jsonl",
|
| 333 |
+
"sha256": "7541893d2b2ec0ec2ffc513f50c568eaef860b65bcef9a1747996c116fff9661"
|
| 334 |
},
|
| 335 |
"v4": {
|
| 336 |
+
"path": "eval/heldout/v4/baselines/base-4b/core/normalized.jsonl",
|
| 337 |
+
"sha256": "7a6d252ccaa72bf54b94445190bbad023fb465a3eaf4cf657532e2cdd4f750f5"
|
| 338 |
},
|
| 339 |
"v5": {
|
| 340 |
+
"path": "eval/heldout/v5/open-baselines-v1/base-4b/core/normalized.jsonl",
|
| 341 |
+
"sha256": "c01ad5daed65267e9aec40ed9e97ece1b6511e4cab17ebff09339b41524fb7f1"
|
| 342 |
}
|
| 343 |
},
|
| 344 |
"transfer_rows": {
|
| 345 |
+
"path": "analysis/decision-benchmark-v3-baselines/base-4b/ROWS.json",
|
| 346 |
+
"sha256": "37ab17d94d640ee46283e1f23723b9ea05133c22a8a9b62a97c756a89c5de785"
|
| 347 |
},
|
| 348 |
"evidence": [
|
| 349 |
{
|
| 350 |
+
"path": "results/remaining-transfer-baselines-v1/base-4b/worker/COMPLETE.json",
|
| 351 |
+
"sha256": "72abc4f87a9234f95eb3b0a72a9e7d66798779d5982c757faffb610003b1b242"
|
| 352 |
},
|
| 353 |
{
|
| 354 |
+
"path": "results/remaining-transfer-baselines-v1/base-4b/worker/METADATA.json",
|
| 355 |
+
"sha256": "871fed92cf762187015e59174a042808dbf3fc88f7791aa0ddd530ae923196ab"
|
| 356 |
},
|
| 357 |
{
|
| 358 |
+
"path": "results/remaining-transfer-baselines-v1/base-4b/worker/predictions.jsonl",
|
| 359 |
+
"sha256": "8c47edda469f8914eccc2ecafa91c245752812e830cb9f29445653102522854a"
|
| 360 |
},
|
| 361 |
{
|
| 362 |
+
"path": "analysis/decision-benchmark-v3-baselines/base-4b/REPORT.json",
|
| 363 |
+
"sha256": "c2a0aafbe2b7c85b3c07a4b9843182a3361654f57883311155aa3b05cfc7002d"
|
| 364 |
},
|
| 365 |
{
|
| 366 |
"path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
|
| 367 |
"sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
|
| 368 |
},
|
| 369 |
{
|
| 370 |
+
"path": "analysis/decision-benchmark-v3-baselines/base-4b/COMPOSABLE-MANIFEST.json",
|
| 371 |
+
"sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
|
| 372 |
+
}
|
| 373 |
+
]
|
| 374 |
+
},
|
| 375 |
+
"Sol": {
|
| 376 |
+
"panels": {
|
| 377 |
+
"old_core": {
|
| 378 |
+
"path": "results/expanded/sol-composition-v2/candidate/job-0/worker/old_core/normalized.jsonl",
|
| 379 |
+
"sha256": "5fc0b3e49a5de576f7621839130855c178b12d942dc357aaf6b4dd9735e0a6a8"
|
| 380 |
+
},
|
| 381 |
+
"v3_core": {
|
| 382 |
+
"path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v3_core/normalized.jsonl",
|
| 383 |
+
"sha256": "b2c99fbd739b849a66a55929a658de4517a7ffaa4496255566e6972ba28d4f15"
|
| 384 |
+
},
|
| 385 |
+
"v4": {
|
| 386 |
+
"path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v4/normalized.jsonl",
|
| 387 |
+
"sha256": "3c28302750e02d4cead0aacd66a323734e773f726c4c41db658dd6461251b28b"
|
| 388 |
+
},
|
| 389 |
+
"v5": {
|
| 390 |
+
"path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v5/normalized.jsonl",
|
| 391 |
+
"sha256": "6ddb0496012b596ea122afa4db4c4158e57dfb934e9ae29eb006756b652e31e1"
|
| 392 |
+
}
|
| 393 |
+
},
|
| 394 |
+
"transfer_rows": {
|
| 395 |
+
"path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/v9-transfer-v9-test-published-temperature-rows.json",
|
| 396 |
+
"sha256": "b05799417cb3715de633135c58f713960b75bdb9ed30eb8912a197d33ee3d85b"
|
| 397 |
+
},
|
| 398 |
+
"evidence": [
|
| 399 |
+
{
|
| 400 |
+
"path": "results/expanded/sol-composition-v2/QUALITY-COLLECTION.json",
|
| 401 |
+
"sha256": "9b77aac87b85cbf3a74e86af4a2b9a03fd9e1ed854356a92ff54ccd0d47a16f4"
|
| 402 |
+
},
|
| 403 |
+
{
|
| 404 |
+
"path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/REPORT.json",
|
| 405 |
+
"sha256": "5d8ebb339293687cec9089f54fb96aabf3330f16111487a49489f5e52bbb6f29"
|
| 406 |
+
}
|
| 407 |
+
]
|
| 408 |
+
},
|
| 409 |
+
"kev-0.8b": {
|
| 410 |
+
"panels": {
|
| 411 |
+
"old_core": {
|
| 412 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-old_core-normalized.jsonl",
|
| 413 |
+
"sha256": "b492b4512ac164f83a319f31bbc9e086cae12f87ffe2c14f307481ffb240807c"
|
| 414 |
+
},
|
| 415 |
+
"v3_core": {
|
| 416 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v3_core-normalized.jsonl",
|
| 417 |
+
"sha256": "70cbfc69dfc72ac0c5ae01305ec63a9f31129c115b7068c6824cd22ead013d68"
|
| 418 |
+
},
|
| 419 |
+
"v4": {
|
| 420 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v4-normalized.jsonl",
|
| 421 |
+
"sha256": "761dfd69acbface101f2c21fd27d2db684c5c8b8cc425890a075729ae1f7db24"
|
| 422 |
+
},
|
| 423 |
+
"v5": {
|
| 424 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v5-normalized.jsonl",
|
| 425 |
+
"sha256": "51b6fc5f826ca8db44df43885f34d857ddd5b5a66453bbdc1a83fcf488928bdb"
|
| 426 |
+
}
|
| 427 |
+
},
|
| 428 |
+
"transfer_rows": {
|
| 429 |
+
"path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/v9-transfer-v9-test-published-temperature-rows.json",
|
| 430 |
+
"sha256": "d6abed88001aeac630219466f02c1aad89fcbd354b3ae533507f3c2f7fe42898"
|
| 431 |
+
},
|
| 432 |
+
"evidence": [
|
| 433 |
+
{
|
| 434 |
+
"path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/REPORT.json",
|
| 435 |
+
"sha256": "6760dcbaa7d85e6c2e893e4c80ab8e66e0b84d7cfa26ae9e25909f48a05f8de4"
|
| 436 |
+
},
|
| 437 |
+
{
|
| 438 |
+
"path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
|
| 439 |
+
"sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
|
| 440 |
}
|
| 441 |
]
|
| 442 |
},
|
|
|
|
| 490 |
}
|
| 491 |
]
|
| 492 |
},
|
| 493 |
+
"Laya-base": {
|
| 494 |
"panels": {
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 495 |
"v4": {
|
| 496 |
+
"path": "results/public-roster-baselines-v1/laya-english/worker/v4/normalized.jsonl",
|
| 497 |
+
"sha256": "31da1d4140da69d2677172707eb301843cf8fdfb43200cba93be910321f922ba"
|
| 498 |
},
|
| 499 |
"v5": {
|
| 500 |
+
"path": "results/public-roster-baselines-v1/laya-english/worker/v5/normalized.jsonl",
|
| 501 |
+
"sha256": "159614ceede814b578e1e025836a595c8c71a9bec4896caf6454d7bf54c4c44d"
|
| 502 |
+
},
|
| 503 |
+
"old_core": {
|
| 504 |
+
"path": "eval/heldout/primary-v2/laya-english/core/normalized.jsonl",
|
| 505 |
+
"sha256": "850643d2daeb47a6909985003736cab0f1f0e5a87631a35e27192f0733efd756"
|
| 506 |
+
},
|
| 507 |
+
"v3_core": {
|
| 508 |
+
"path": "eval/heldout/v3/baselines/laya-english/core/normalized.jsonl",
|
| 509 |
+
"sha256": "dec19caeb8074e25289e43a367392172da21f0c2ccdb403a059e2e4a2364a64b"
|
| 510 |
}
|
| 511 |
},
|
| 512 |
"transfer_rows": {
|
| 513 |
+
"path": "analysis/decision-benchmark-v3-baselines/laya-english/ROWS.json",
|
| 514 |
+
"sha256": "19138909259400facb5d3d5402d0d736f870695fe3ee46019f2e438088c57275"
|
| 515 |
},
|
| 516 |
"evidence": [
|
| 517 |
{
|
| 518 |
+
"path": "results/public-roster-baselines-v1/laya-english/worker/COMPLETE.json",
|
| 519 |
+
"sha256": "cef0da797ae14cebaf1e4b9edf9b16c3cd9a412a30435d7d2517b8ba6dbe8b98"
|
| 520 |
},
|
| 521 |
{
|
| 522 |
+
"path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
|
| 523 |
+
"sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
|
| 524 |
},
|
| 525 |
{
|
| 526 |
+
"path": "analysis/decision-benchmark-v3-baselines/laya-english/REPORT.json",
|
| 527 |
+
"sha256": "cb5935bb399dcf23ffe24f855452b2e074bcbb3d023cfc7dd439342e54f02c6b"
|
| 528 |
},
|
| 529 |
{
|
| 530 |
+
"path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
|
| 531 |
+
"sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
|
| 532 |
+
}
|
| 533 |
+
]
|
| 534 |
+
},
|
| 535 |
+
"Laya-multilingual": {
|
| 536 |
+
"panels": {
|
| 537 |
+
"v4": {
|
| 538 |
+
"path": "results/public-roster-baselines-v1/laya-multilingual/worker/v4/normalized.jsonl",
|
| 539 |
+
"sha256": "814c1aa1f4d400942984ecb3786d3fd5fd15286d9b767a519c91993577795290"
|
| 540 |
+
},
|
| 541 |
+
"v5": {
|
| 542 |
+
"path": "results/public-roster-baselines-v1/laya-multilingual/worker/v5/normalized.jsonl",
|
| 543 |
+
"sha256": "6d4dbc16c7b977b15895e0e313dabc6d9db43df75ec56e89373d91f4f13ec676"
|
| 544 |
+
},
|
| 545 |
+
"old_core": {
|
| 546 |
+
"path": "eval/heldout/primary-v2/laya-multilingual/core/normalized.jsonl",
|
| 547 |
+
"sha256": "9ea39841c4d5901cd9ab860c980d2f3b55358be289c7e01b9b47982927bf849d"
|
| 548 |
},
|
| 549 |
+
"v3_core": {
|
| 550 |
+
"path": "eval/heldout/v3/baselines/laya-multilingual/core/normalized.jsonl",
|
| 551 |
+
"sha256": "a0b7ce9daf4be2c7831bcab6b5e96c8dcb6bef133d39475f72241ab18faccdd2"
|
| 552 |
+
}
|
| 553 |
+
},
|
| 554 |
+
"transfer_rows": {
|
| 555 |
+
"path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/ROWS.json",
|
| 556 |
+
"sha256": "d32db6cf6576f287210d31336456497f9d6926179366db29304659baf9783b43"
|
| 557 |
+
},
|
| 558 |
+
"evidence": [
|
| 559 |
{
|
| 560 |
+
"path": "results/public-roster-baselines-v1/laya-multilingual/worker/COMPLETE.json",
|
| 561 |
+
"sha256": "b190d683fc4c4762c2805a03d52628a53e125d228a3105c8ac2ddda00a5364e4"
|
| 562 |
},
|
| 563 |
{
|
| 564 |
+
"path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
|
| 565 |
+
"sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
|
| 566 |
+
},
|
| 567 |
+
{
|
| 568 |
+
"path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/REPORT.json",
|
| 569 |
+
"sha256": "d20d3d58f76ca2758545f83d358abd747f86d581c897d1fb521371d1000d6154"
|
| 570 |
+
},
|
| 571 |
+
{
|
| 572 |
+
"path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
|
| 573 |
+
"sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
|
| 574 |
}
|
| 575 |
]
|
| 576 |
}
|
| 577 |
},
|
| 578 |
"task_rows": 54,
|
| 579 |
"rank_public_models": [
|
|
|
|
| 580 |
"Lux",
|
| 581 |
"Nox",
|
| 582 |
"kev-9b",
|
|
|
|
| 588 |
"kev-0.8b",
|
| 589 |
"Qwen3.5-2B",
|
| 590 |
"Laya-base",
|
| 591 |
+
"Laya-multilingual",
|
| 592 |
+
"Jev"
|
| 593 |
],
|
| 594 |
"internal_extra_models_not_public_rank": [
|
| 595 |
"llm2jev-2b",
|
|
|
|
| 672 |
}
|
| 673 |
},
|
| 674 |
"source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
|
| 675 |
+
}
|
|
|
|
|
|
|
| 676 |
}
|
model-card-example.json
CHANGED
|
@@ -65,7 +65,7 @@
|
|
| 65 |
"direct_engine_exact_response": true,
|
| 66 |
"overflow_rejected": true,
|
| 67 |
"overflow_message": "check: 40083 tokens exceeds max_length=16384; no truncation allowed",
|
| 68 |
-
"bundle_manifest_sha256": "
|
| 69 |
"runtime": {
|
| 70 |
"actual": {
|
| 71 |
"torch": "2.12.0+git6bbd260",
|
|
|
|
| 65 |
"direct_engine_exact_response": true,
|
| 66 |
"overflow_rejected": true,
|
| 67 |
"overflow_message": "check: 40083 tokens exceeds max_length=16384; no truncation allowed",
|
| 68 |
+
"bundle_manifest_sha256": "7d9b06bc25a75b1f131aabd280df0ef2777bf69db98574a300067ace40c61c9c",
|
| 69 |
"runtime": {
|
| 70 |
"actual": {
|
| 71 |
"torch": "2.12.0+git6bbd260",
|
release-manifest.json
CHANGED
|
@@ -1,12 +1,12 @@
|
|
| 1 |
{
|
| 2 |
"format": "decision-public-release-v1",
|
| 3 |
-
"status": "
|
| 4 |
-
"bundle_manifest_sha256": "
|
| 5 |
"readiness_sha256": "fcbb84a9bf7294ac79e2421cdd2cbfdfc46a79fd74c41ba0965778472d1cafc1",
|
| 6 |
-
"model_card_sha256": "
|
| 7 |
"repo_id": "llm-semantic-router/Decision-1.0-Nox",
|
| 8 |
-
"assembly_script_sha256": "
|
| 9 |
-
"original_bundle_manifest_preserved":
|
| 10 |
"files_exclude_this_manifest": true,
|
| 11 |
"files": [
|
| 12 |
{
|
|
@@ -17,7 +17,7 @@
|
|
| 17 |
{
|
| 18 |
"file": "DIAGNOSTICS.md",
|
| 19 |
"bytes": 5678,
|
| 20 |
-
"sha256": "
|
| 21 |
},
|
| 22 |
{
|
| 23 |
"file": "Dockerfile.runtime",
|
|
@@ -26,8 +26,8 @@
|
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"file": "EVALUATION.md",
|
| 29 |
-
"bytes":
|
| 30 |
-
"sha256": "
|
| 31 |
},
|
| 32 |
{
|
| 33 |
"file": "LICENSE",
|
|
@@ -36,14 +36,19 @@
|
|
| 36 |
},
|
| 37 |
{
|
| 38 |
"file": "MATERIALS.json",
|
| 39 |
-
"bytes":
|
| 40 |
-
"sha256": "
|
| 41 |
},
|
| 42 |
{
|
| 43 |
"file": "NORMALIZATION_RUNTIME.md",
|
| 44 |
"bytes": 756,
|
| 45 |
"sha256": "cf8e6ce1f07687a68b6adeb98e6704b8f292cb6adc7b10e4b84e8e70aa5472e5"
|
| 46 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 47 |
{
|
| 48 |
"file": "QUESTION-SCALING.md",
|
| 49 |
"bytes": 1274,
|
|
@@ -56,13 +61,13 @@
|
|
| 56 |
},
|
| 57 |
{
|
| 58 |
"file": "README.md",
|
| 59 |
-
"bytes":
|
| 60 |
-
"sha256": "
|
| 61 |
},
|
| 62 |
{
|
| 63 |
"file": "RUNTIME-RELEASE.json",
|
| 64 |
-
"bytes":
|
| 65 |
-
"sha256": "
|
| 66 |
},
|
| 67 |
{
|
| 68 |
"file": "RUNTIME.md",
|
|
@@ -76,8 +81,8 @@
|
|
| 76 |
},
|
| 77 |
{
|
| 78 |
"file": "SENSITIVITY.md",
|
| 79 |
-
"bytes":
|
| 80 |
-
"sha256": "
|
| 81 |
},
|
| 82 |
{
|
| 83 |
"file": "SERVING_OPTIMIZATION.json",
|
|
@@ -91,13 +96,18 @@
|
|
| 91 |
},
|
| 92 |
{
|
| 93 |
"file": "TASKS.md",
|
| 94 |
-
"bytes":
|
| 95 |
-
"sha256": "
|
| 96 |
},
|
| 97 |
{
|
| 98 |
"file": "USAGE.md",
|
| 99 |
-
"bytes":
|
| 100 |
-
"sha256": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 101 |
},
|
| 102 |
{
|
| 103 |
"file": "assets/architecture.png",
|
|
@@ -236,18 +246,18 @@
|
|
| 236 |
},
|
| 237 |
{
|
| 238 |
"file": "assets/decision-matrix.pdf",
|
| 239 |
-
"bytes":
|
| 240 |
-
"sha256": "
|
| 241 |
},
|
| 242 |
{
|
| 243 |
"file": "assets/decision-matrix.png",
|
| 244 |
-
"bytes":
|
| 245 |
-
"sha256": "
|
| 246 |
},
|
| 247 |
{
|
| 248 |
"file": "assets/decision-matrix.svg",
|
| 249 |
"bytes": 43846,
|
| 250 |
-
"sha256": "
|
| 251 |
},
|
| 252 |
{
|
| 253 |
"file": "assets/decision-question-scaling-600px.png",
|
|
@@ -271,18 +281,18 @@
|
|
| 271 |
},
|
| 272 |
{
|
| 273 |
"file": "assets/decision-ranking.pdf",
|
| 274 |
-
"bytes":
|
| 275 |
-
"sha256": "
|
| 276 |
},
|
| 277 |
{
|
| 278 |
"file": "assets/decision-ranking.png",
|
| 279 |
-
"bytes":
|
| 280 |
-
"sha256": "
|
| 281 |
},
|
| 282 |
{
|
| 283 |
"file": "assets/decision-ranking.svg",
|
| 284 |
"bytes": 14180,
|
| 285 |
-
"sha256": "
|
| 286 |
},
|
| 287 |
{
|
| 288 |
"file": "assets/readout.png",
|
|
@@ -316,8 +326,8 @@
|
|
| 316 |
},
|
| 317 |
{
|
| 318 |
"file": "bundle-manifest.json",
|
| 319 |
-
"bytes":
|
| 320 |
-
"sha256": "
|
| 321 |
},
|
| 322 |
{
|
| 323 |
"file": "chat_template.jinja",
|
|
@@ -326,8 +336,8 @@
|
|
| 326 |
},
|
| 327 |
{
|
| 328 |
"file": "code/decision_api.py",
|
| 329 |
-
"bytes":
|
| 330 |
-
"sha256": "
|
| 331 |
},
|
| 332 |
{
|
| 333 |
"file": "code/decision_model.py",
|
|
@@ -361,8 +371,8 @@
|
|
| 361 |
},
|
| 362 |
{
|
| 363 |
"file": "metrics/benchmark.json",
|
| 364 |
-
"bytes":
|
| 365 |
-
"sha256": "
|
| 366 |
},
|
| 367 |
{
|
| 368 |
"file": "metrics/comparator-coverage.json",
|
|
@@ -371,8 +381,8 @@
|
|
| 371 |
},
|
| 372 |
{
|
| 373 |
"file": "metrics/evaluation-provenance.json",
|
| 374 |
-
"bytes":
|
| 375 |
-
"sha256": "
|
| 376 |
},
|
| 377 |
{
|
| 378 |
"file": "metrics/expanded-quality.json",
|
|
@@ -407,7 +417,7 @@
|
|
| 407 |
{
|
| 408 |
"file": "model-card-example.json",
|
| 409 |
"bytes": 3775,
|
| 410 |
-
"sha256": "
|
| 411 |
},
|
| 412 |
{
|
| 413 |
"file": "pyproject.toml",
|
|
@@ -557,10 +567,10 @@
|
|
| 557 |
"MATERIALS.json": "copy",
|
| 558 |
"metrics/semantic-consistency.json": "copy"
|
| 559 |
},
|
| 560 |
-
"scope": "
|
| 561 |
-
"release_tag": "v1.3.
|
| 562 |
-
"change_kind": "
|
| 563 |
-
"previous_main_revision": "
|
| 564 |
-
"previous_release_manifest_sha256": "
|
| 565 |
"presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
|
| 566 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"format": "decision-public-release-v1",
|
| 3 |
+
"status": "qualified-runtime-and-current-documents-assembled",
|
| 4 |
+
"bundle_manifest_sha256": "7d9b06bc25a75b1f131aabd280df0ef2777bf69db98574a300067ace40c61c9c",
|
| 5 |
"readiness_sha256": "fcbb84a9bf7294ac79e2421cdd2cbfdfc46a79fd74c41ba0965778472d1cafc1",
|
| 6 |
+
"model_card_sha256": "41dfa047d33bdc3aa9feff55928972533ba803a30b332ca380f042268fa1e848",
|
| 7 |
"repo_id": "llm-semantic-router/Decision-1.0-Nox",
|
| 8 |
+
"assembly_script_sha256": "a8f72b8a67a64034d6183f89c945bab611056816ba0cba43670e9becb6eb78c4",
|
| 9 |
+
"original_bundle_manifest_preserved": false,
|
| 10 |
"files_exclude_this_manifest": true,
|
| 11 |
"files": [
|
| 12 |
{
|
|
|
|
| 17 |
{
|
| 18 |
"file": "DIAGNOSTICS.md",
|
| 19 |
"bytes": 5678,
|
| 20 |
+
"sha256": "0b290793565cc5007e9ee7bc79625812c73202c7d8bfce79d8816a0e7a93a965"
|
| 21 |
},
|
| 22 |
{
|
| 23 |
"file": "Dockerfile.runtime",
|
|
|
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"file": "EVALUATION.md",
|
| 29 |
+
"bytes": 3705,
|
| 30 |
+
"sha256": "4159b7939a107e91b13e47345a192eea489f8bd870943e4b397e094b33bba905"
|
| 31 |
},
|
| 32 |
{
|
| 33 |
"file": "LICENSE",
|
|
|
|
| 36 |
},
|
| 37 |
{
|
| 38 |
"file": "MATERIALS.json",
|
| 39 |
+
"bytes": 2111,
|
| 40 |
+
"sha256": "6591d0318677c51084db78bd3bfc09bb5cacaa6c9c2d5700fc969f64b9a7ac4c"
|
| 41 |
},
|
| 42 |
{
|
| 43 |
"file": "NORMALIZATION_RUNTIME.md",
|
| 44 |
"bytes": 756,
|
| 45 |
"sha256": "cf8e6ce1f07687a68b6adeb98e6704b8f292cb6adc7b10e4b84e8e70aa5472e5"
|
| 46 |
},
|
| 47 |
+
{
|
| 48 |
+
"file": "NULL_DESCRIPTION_RENDERING.json",
|
| 49 |
+
"bytes": 514,
|
| 50 |
+
"sha256": "3b531cab60cba35648ab71fe4be7fd1af619f02ee337e3beed6296fdfc8b9ec8"
|
| 51 |
+
},
|
| 52 |
{
|
| 53 |
"file": "QUESTION-SCALING.md",
|
| 54 |
"bytes": 1274,
|
|
|
|
| 61 |
},
|
| 62 |
{
|
| 63 |
"file": "README.md",
|
| 64 |
+
"bytes": 4767,
|
| 65 |
+
"sha256": "41dfa047d33bdc3aa9feff55928972533ba803a30b332ca380f042268fa1e848"
|
| 66 |
},
|
| 67 |
{
|
| 68 |
"file": "RUNTIME-RELEASE.json",
|
| 69 |
+
"bytes": 679,
|
| 70 |
+
"sha256": "325d7e037acab5f61dab4d56d817e7d7e2e873d9b5f0d0df587afd6c33364a46"
|
| 71 |
},
|
| 72 |
{
|
| 73 |
"file": "RUNTIME.md",
|
|
|
|
| 81 |
},
|
| 82 |
{
|
| 83 |
"file": "SENSITIVITY.md",
|
| 84 |
+
"bytes": 788,
|
| 85 |
+
"sha256": "c3ea31420db4f2e5f17cc7197197a3ddd18abdad773ffb5bf1171b67136c3bc2"
|
| 86 |
},
|
| 87 |
{
|
| 88 |
"file": "SERVING_OPTIMIZATION.json",
|
|
|
|
| 96 |
},
|
| 97 |
{
|
| 98 |
"file": "TASKS.md",
|
| 99 |
+
"bytes": 9299,
|
| 100 |
+
"sha256": "2fb04bc9c17d805e2abd5e90520df25e79d85aa5296cd42074d4b156f54bf862"
|
| 101 |
},
|
| 102 |
{
|
| 103 |
"file": "USAGE.md",
|
| 104 |
+
"bytes": 5665,
|
| 105 |
+
"sha256": "b4c2f44ce120149c7cdee3749283958f34ba9caf4b6492442818c1b6a748c0d3"
|
| 106 |
+
},
|
| 107 |
+
{
|
| 108 |
+
"file": "WEIGHTING.md",
|
| 109 |
+
"bytes": 788,
|
| 110 |
+
"sha256": "c3ea31420db4f2e5f17cc7197197a3ddd18abdad773ffb5bf1171b67136c3bc2"
|
| 111 |
},
|
| 112 |
{
|
| 113 |
"file": "assets/architecture.png",
|
|
|
|
| 246 |
},
|
| 247 |
{
|
| 248 |
"file": "assets/decision-matrix.pdf",
|
| 249 |
+
"bytes": 28045,
|
| 250 |
+
"sha256": "5e127c0cd1a6f83b772a3e66ac2e1f4ec19913ac4938371c6f190a23c50a29fe"
|
| 251 |
},
|
| 252 |
{
|
| 253 |
"file": "assets/decision-matrix.png",
|
| 254 |
+
"bytes": 325781,
|
| 255 |
+
"sha256": "c775737aedeecadb0824092d37014a8e2cffa90af0dafcf10d584e0db823ecf8"
|
| 256 |
},
|
| 257 |
{
|
| 258 |
"file": "assets/decision-matrix.svg",
|
| 259 |
"bytes": 43846,
|
| 260 |
+
"sha256": "d99102fafe842f2a6bdd4d455f915833d5e0e09e3a22e5211fbf86e5647247ab"
|
| 261 |
},
|
| 262 |
{
|
| 263 |
"file": "assets/decision-question-scaling-600px.png",
|
|
|
|
| 281 |
},
|
| 282 |
{
|
| 283 |
"file": "assets/decision-ranking.pdf",
|
| 284 |
+
"bytes": 24204,
|
| 285 |
+
"sha256": "a28d4b0e8649cb7bf2db452a11d93675631af795561726aa63432b05b5888049"
|
| 286 |
},
|
| 287 |
{
|
| 288 |
"file": "assets/decision-ranking.png",
|
| 289 |
+
"bytes": 215215,
|
| 290 |
+
"sha256": "d252593501e8faeb9c5c8cc0a9ae9a96591e7aa8696cf5394cdfec9c0ddac3b7"
|
| 291 |
},
|
| 292 |
{
|
| 293 |
"file": "assets/decision-ranking.svg",
|
| 294 |
"bytes": 14180,
|
| 295 |
+
"sha256": "ab7ca66a9202587dee593714e46be760cb12ac91b6459074bcbbe11c59107d9d"
|
| 296 |
},
|
| 297 |
{
|
| 298 |
"file": "assets/readout.png",
|
|
|
|
| 326 |
},
|
| 327 |
{
|
| 328 |
"file": "bundle-manifest.json",
|
| 329 |
+
"bytes": 8171,
|
| 330 |
+
"sha256": "7d9b06bc25a75b1f131aabd280df0ef2777bf69db98574a300067ace40c61c9c"
|
| 331 |
},
|
| 332 |
{
|
| 333 |
"file": "chat_template.jinja",
|
|
|
|
| 336 |
},
|
| 337 |
{
|
| 338 |
"file": "code/decision_api.py",
|
| 339 |
+
"bytes": 10978,
|
| 340 |
+
"sha256": "6716bdabca3cd2f1aa447d42a6cf53ee62d7ea97984a01f16b4c09fac8aaf17e"
|
| 341 |
},
|
| 342 |
{
|
| 343 |
"file": "code/decision_model.py",
|
|
|
|
| 371 |
},
|
| 372 |
{
|
| 373 |
"file": "metrics/benchmark.json",
|
| 374 |
+
"bytes": 2261123,
|
| 375 |
+
"sha256": "178e2e5f58da45f9fbf378ca82b49d7878fc8e9c93eee195a5ef1e7ba0f4fc6b"
|
| 376 |
},
|
| 377 |
{
|
| 378 |
"file": "metrics/comparator-coverage.json",
|
|
|
|
| 381 |
},
|
| 382 |
{
|
| 383 |
"file": "metrics/evaluation-provenance.json",
|
| 384 |
+
"bytes": 29254,
|
| 385 |
+
"sha256": "a3718fcee42a8cfe0144b01e04346c36a09ac7958154e4561327753d87d5fe8c"
|
| 386 |
},
|
| 387 |
{
|
| 388 |
"file": "metrics/expanded-quality.json",
|
|
|
|
| 417 |
{
|
| 418 |
"file": "model-card-example.json",
|
| 419 |
"bytes": 3775,
|
| 420 |
+
"sha256": "1c4fc87d543722e2c3b839cbebd30c5d1370fcaf2d897f97de13079a668ac93e"
|
| 421 |
},
|
| 422 |
{
|
| 423 |
"file": "pyproject.toml",
|
|
|
|
| 567 |
"MATERIALS.json": "copy",
|
| 568 |
"metrics/semantic-consistency.json": "copy"
|
| 569 |
},
|
| 570 |
+
"scope": "Latest13model five-panel comparison and54task diagnostics; model tensor/tokenizer/temperature identities unchanged.",
|
| 571 |
+
"release_tag": "v1.3.2",
|
| 572 |
+
"change_kind": "null-Choice-description-runtime-and-current-materials",
|
| 573 |
+
"previous_main_revision": "46505c737a45cbe2c4ac4eee38e4f94cb520e4ed",
|
| 574 |
+
"previous_release_manifest_sha256": "f537b023f52d079fdd77ecb5b1240799959209de172cc6162a6996a009ef7d03",
|
| 575 |
"presentation_amendment_sha256": "f078ecb12e7f4f2baaab522d07af6695c18d9c4669a5f61623e85c93b910fd1f"
|
| 576 |
}
|