thegovind commited on
Commit
57ae92b
·
verified ·
1 Parent(s): 54ec0c5

Card: local Decision Index 0.2 run (descriptive; known training exposure not penalized) and its evidence file

Browse files
Files changed (2) hide show
  1. README.md +51 -4
  2. eval/decision-index-0.2-local.json +1482 -0
README.md CHANGED
@@ -32,6 +32,55 @@ Xiaomi, Alibaba Cloud or the Qwen team. Weights are for non-commercial research;
32
 
33
  ## Results
34
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
35
  ### Decision Index 0.1 (archived edition)
36
 
37
  **Full suite: blink-mimo-9b 56.53 vs Jev 1.13.0 59.51.**
@@ -48,9 +97,7 @@ Xiaomi, Alibaba Cloud or the Qwen team. Weights are for non-commercial research;
48
  | Kev 9B | 9B | 50.48 | 32.96 | 30.54 |
49
  | Kev 4B | 4B | 47.43 | 28.86 | 25.67 |
50
 
51
- We ran the complete archived 0.1 suite: 132,422 requests across 37 benchmarks. The headline index averages 19 panel benchmarks. Comparison rows use the 2026-09-22 leaderboard snapshot. We ran the official kit's scorer locally; these aren't leaderboard submissions. The live [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index) moved to 0.2 on 2026-09-24, but the public kit can't build 0.2 yet.
52
-
53
- No Decision Index 0.2 result is reported for these models. Comparable shared-benchmark results require matched request subsets and the 0.2 metric transformations.
54
 
55
  Third in our local archived 0.1 comparison, behind blink-27b and Jev and ahead of every open entry in the September 22 snapshot (best: Jevfire).
56
 
@@ -287,7 +334,7 @@ These are source-repository licences; they don't settle rights in every underlyi
287
  - **Training overlap.** Public train splits also used by the 0.1 index: ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, ChessBench (searchless_chess training positions; none of the 5,000 test positions), GSM8K (train split; solution-checking items). We also used ANLI and BANKING77 train splits; they're in the 0.1 suite but outside its index, and both are in the 0.2 panel. The audit below reports what was checked and any shared passages.
288
  - **Partitions.** Public-source data included training and development partitions.
289
  - **Final-mixture audit.** Rechecked every question row (including teacher-written rows) against the complete 0.1 suite (132,422 requests) and JevBench's 231 public items. The checks looked for exact matches of normalised strings of at least 30 characters in any field and shared 13-word passages in each row's question text (instructions, state.question, state.code). Strings or passages seen in 20 or more suite requests were treated as prompt templates and ignored. No public JevBench item matched under these checks; a separate position check found no shared chess positions.
290
- - **Suite overlap.** 16 BANKING77/VAST training rows share a 13-word passage with 31 suite requests: 23 of VAST's 3,006 and 8 of BANKING77's 3,080. Two VAST training posts are near-duplicates of a test post, but these rows had no exact normalised-text match under the audit. Dropping those requests leaves the index at 56.53 (VAST 0.7805 → 0.7803); BANKING77 is outside the index.
291
  - **Audit limits.** The 13-word passage check didn't search long-document bodies or option text. Semantic or pretraining overlap can't be ruled out, and private JevBench items weren't available to check.
292
  - **Generated reasoning.** Our programs computed the labels for CRUXEval-style code and CLadder-style causal questions; no items from those benchmarks were used. We didn't reuse the suite's GSM8K distractors.
293
  - **Teacher documents.** We kept Qwen3.8-27B's documents only if a fresh blind solve by that same teacher agreed with the answer. That's an agreement filter, not independent verification.
 
32
 
33
  ## Results
34
 
35
+ ### Decision Index 0.2 (local run)
36
+
37
+ We ran the full Decision Index 0.2 suite ourselves with the official scoring kit at commit 19ad28e on 2026-09-25. This is a descriptive run, not a leaderboard submission or accepted result. The kit's scorer does not apply the leaderboard's penalty for rows an entrant trained on, so known training exposure stays in these scores and they cannot be ranked against the leaderboard.
38
+
39
+ | Balanced skill | Balanced raw | Breadth skill | Without MMLU-Pro |
40
+ |---:|---:|---:|---:|
41
+ | 43.36 | 57.24 | 42.38 | 42.84 |
42
+
43
+ - Our Blink fine-tune did not use MMLU-Pro or GPQA as direct data sources. Screening found 176 SuperGPQA training rows matching added-request text. SuperGPQA is a training source, not a benchmark in this suite. Public train splits used in training are listed under “Training overlap” below.
44
+
45
+ <details><summary>Extra tables and method</summary>
46
+
47
+ | Area | Number of benchmarks | Skill | Raw |
48
+ |---|---:|---:|---:|
49
+ | Knowledge & Reasoning | 10 | 33.2 | 48.2 |
50
+ | Language Understanding | 10 | 54.1 | 67.2 |
51
+ | Retrieval & Classification | 7 | 40.5 | 55.4 |
52
+ | Tools & Automation | 6 | 57.0 | 65.2 |
53
+ | Arts & Human Taste | 7 | 32.0 | 50.2 |
54
+
55
+ The seven benchmarks added in 0.2.
56
+
57
+ | Benchmark | Metric | Requests | Answered | Raw | Skill |
58
+ |---|---|---:|---:|---:|---:|
59
+ | PhishNChips phishing decisions | accuracy | 2,000 | 2,000 | 66.5 | 33.1 |
60
+ | MMLU-Pro | accuracy | 12,032 | 12,032 | 61.4 | 56.5 |
61
+ | BBH fixed-option tasks | accuracy | 5,507 | 5,507 | 67.8 | 53.4 |
62
+ | RAGTruth response-level hallucination | F1 on hallucinated class | 2,700 | 2,700 | 63.4 | 37.8 |
63
+ | HoVer claim verification | accuracy | 4,000 | 4,000 | 65.6 | 31.3 |
64
+ | When2Call MCQ | accuracy | 3,652 | 3,652 | 64.5 | 52.7 |
65
+ | New Yorker caption matching | accuracy | 528 | 528 | 62.5 | 53.1 |
66
+
67
+ - All 151,034 of 151,034 scoreable requests scored. Of 44 scored benchmarks, 40 count toward the index across five equal areas.
68
+ - Requests shared with 0.1 reuse the model's 0.1 predictions. We ran the 30,419 added requests with the same frozen evaluation setup as 0.1, the evaluated adapter loaded on the MiMo base, which the published graft matched on JevBench's 231 public items (see Evaluation notes), at temperature 1.0.
69
+ - Balanced skill is the headline index. “Without MMLU-Pro” drops MMLU-Pro, averages the other nine Knowledge benchmarks, and keeps five equal areas. It is a sensitivity check, not a score free of training effects.
70
+ - These are point estimates, with no significance, calibration, or latency claims. Do not compare them with 0.1 numbers because the editions differ.
71
+
72
+ Training-row text matches in the added requests.
73
+
74
+ | Training stage | Rows in the stage | Rows matching added-request text | From MMLU-Pro | From SuperGPQA | Other |
75
+ |---|---:|---:|---:|---:|---:|
76
+ | MiMo | 123,195 | 180 | 0 | 176 | 4 |
77
+
78
+ We screened for exact normalised strings of at least 30 characters shared by training rows and added requests, ignoring strings found in 20 or more requests as templates. Counts are training rows by stage and source, not unique test questions. A matching option or passage need not be the same question, and a clean screen cannot rule out semantic or pretraining overlap.
79
+
80
+ We did not produce the planned calibration read or a score without the DI-S selection sample.
81
+
82
+ </details>
83
+
84
  ### Decision Index 0.1 (archived edition)
85
 
86
  **Full suite: blink-mimo-9b 56.53 vs Jev 1.13.0 59.51.**
 
97
  | Kev 9B | 9B | 50.48 | 32.96 | 30.54 |
98
  | Kev 4B | 4B | 47.43 | 28.86 | 25.67 |
99
 
100
+ We ran the complete archived 0.1 suite: 132,422 requests across 37 benchmarks. The headline index averages 19 panel benchmarks. Comparison rows use the 2026-09-22 leaderboard snapshot. We ran the official kit's scorer locally; these aren't leaderboard submissions. The live [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index) moved to 0.2 on 2026-09-24. Our local 0.2 run is in the section above.
 
 
101
 
102
  Third in our local archived 0.1 comparison, behind blink-27b and Jev and ahead of every open entry in the September 22 snapshot (best: Jevfire).
103
 
 
334
  - **Training overlap.** Public train splits also used by the 0.1 index: ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, ChessBench (searchless_chess training positions; none of the 5,000 test positions), GSM8K (train split; solution-checking items). We also used ANLI and BANKING77 train splits; they're in the 0.1 suite but outside its index, and both are in the 0.2 panel. The audit below reports what was checked and any shared passages.
335
  - **Partitions.** Public-source data included training and development partitions.
336
  - **Final-mixture audit.** Rechecked every question row (including teacher-written rows) against the complete 0.1 suite (132,422 requests) and JevBench's 231 public items. The checks looked for exact matches of normalised strings of at least 30 characters in any field and shared 13-word passages in each row's question text (instructions, state.question, state.code). Strings or passages seen in 20 or more suite requests were treated as prompt templates and ignored. No public JevBench item matched under these checks; a separate position check found no shared chess positions.
337
+ - **Suite overlap.** 16 BANKING77/VAST training rows share a 13-word passage with 31 suite requests: 23 of VAST's 3,006 and 8 of BANKING77's 3,080. Two VAST training posts are near-duplicates of a test post, but these rows had no exact normalised-text match under the audit. Dropping those requests leaves the index at 56.53 (VAST 0.7805 → 0.7803); BANKING77 is outside the 0.1 index.
338
  - **Audit limits.** The 13-word passage check didn't search long-document bodies or option text. Semantic or pretraining overlap can't be ruled out, and private JevBench items weren't available to check.
339
  - **Generated reasoning.** Our programs computed the labels for CRUXEval-style code and CLadder-style causal questions; no items from those benchmarks were used. We didn't reuse the suite's GSM8K distractors.
340
  - **Teacher documents.** We kept Qwen3.8-27B's documents only if a fresh blind solve by that same teacher agreed with the answer. That's an agreement filter, not independent verification.
eval/decision-index-0.2-local.json ADDED
@@ -0,0 +1,1482 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "engine": "blink-mimo-9b",
3
+ "edition": "Decision Index 0.2 (local run, descriptive)",
4
+ "generated_utc": "2026-09-25T23:10:24+00:00",
5
+ "suite": {
6
+ "edition": "release-v2",
7
+ "requests": 121057,
8
+ "scoreable": 120615,
9
+ "excluded": 442,
10
+ "added_requests": 30419,
11
+ "benchmarks": 44,
12
+ "rows_sha256": "b2b56d6fb636837ca469e689087bdbf373dda8de7638aa2da6793e6eda0792d5",
13
+ "added_sha256": "7429f3c9cdddb772c1cfc42bb2a45e8516b0032152b746e6929f1c8b52f4ce89"
14
+ },
15
+ "completed": 151034,
16
+ "complete": true,
17
+ "counts": {
18
+ "ok": 151034
19
+ },
20
+ "latency_ms": {
21
+ "median": 45.8,
22
+ "p95": 1053.9,
23
+ "mean": 214.9
24
+ },
25
+ "decision_index": 43.36,
26
+ "raw_index": 57.24,
27
+ "scores": {
28
+ "balanced_skill": 43.36,
29
+ "balanced_raw": 57.24,
30
+ "breadth_skill": 42.38
31
+ },
32
+ "areas": [
33
+ {
34
+ "id": "knowledge",
35
+ "label": "Knowledge & Reasoning",
36
+ "raw": 0.4824,
37
+ "skill": 0.3319,
38
+ "coverage": 1.0,
39
+ "n": 10,
40
+ "benchmarks": [
41
+ 25,
42
+ 30,
43
+ 31,
44
+ 32,
45
+ 33,
46
+ 43,
47
+ 44,
48
+ 45,
49
+ 57,
50
+ 58
51
+ ]
52
+ },
53
+ {
54
+ "id": "language",
55
+ "label": "Language Understanding",
56
+ "raw": 0.6719,
57
+ "skill": 0.5405,
58
+ "coverage": 1.0,
59
+ "n": 10,
60
+ "benchmarks": [
61
+ 11,
62
+ 12,
63
+ 28,
64
+ 29,
65
+ 38,
66
+ 39,
67
+ 40,
68
+ 41,
69
+ 42,
70
+ 59
71
+ ]
72
+ },
73
+ {
74
+ "id": "retrieval",
75
+ "label": "Retrieval & Classification",
76
+ "raw": 0.5536,
77
+ "skill": 0.4048,
78
+ "coverage": 1.0,
79
+ "n": 7,
80
+ "benchmarks": [
81
+ 4,
82
+ 5,
83
+ 10,
84
+ 36,
85
+ 37,
86
+ 56,
87
+ 61
88
+ ]
89
+ },
90
+ {
91
+ "id": "tools",
92
+ "label": "Tools & Automation",
93
+ "raw": 0.6519,
94
+ "skill": 0.5702,
95
+ "coverage": 1.0,
96
+ "n": 6,
97
+ "benchmarks": [
98
+ 1,
99
+ 2,
100
+ 3,
101
+ 6,
102
+ 9,
103
+ 62
104
+ ]
105
+ },
106
+ {
107
+ "id": "arts",
108
+ "label": "Arts & Human Taste",
109
+ "raw": 0.502,
110
+ "skill": 0.3204,
111
+ "coverage": 1.0,
112
+ "n": 7,
113
+ "benchmarks": [
114
+ 20,
115
+ 21,
116
+ 22,
117
+ 23,
118
+ 48,
119
+ 50,
120
+ 64
121
+ ]
122
+ }
123
+ ],
124
+ "index_benchmarks": {
125
+ "1": {
126
+ "raw": 0.8926,
127
+ "skill": 0.855,
128
+ "coverage": 1.0,
129
+ "random": 0.2592,
130
+ "rule": "track",
131
+ "in_index": true,
132
+ "tracks": []
133
+ },
134
+ "2": {
135
+ "raw": 0.4296,
136
+ "skill": 0.3719,
137
+ "coverage": 1.0,
138
+ "random": 0.0918,
139
+ "rule": "track",
140
+ "in_index": true,
141
+ "tracks": []
142
+ },
143
+ "3": {
144
+ "raw": 0.8386,
145
+ "skill": 0.8355,
146
+ "coverage": 1.0,
147
+ "random": 0.0189,
148
+ "rule": "chance",
149
+ "in_index": true
150
+ },
151
+ "4": {
152
+ "raw": 0.8177,
153
+ "skill": 0.8154,
154
+ "coverage": 1.0,
155
+ "random": 0.0127,
156
+ "rule": "chance",
157
+ "in_index": true
158
+ },
159
+ "5": {
160
+ "raw": 0.8381,
161
+ "skill": 0.8371,
162
+ "coverage": 1.0,
163
+ "random": 0.006,
164
+ "rule": "chance",
165
+ "in_index": true
166
+ },
167
+ "6": {
168
+ "raw": 0.7989,
169
+ "skill": 0.5256,
170
+ "coverage": 1.0,
171
+ "random": 0.5245,
172
+ "rule": "track",
173
+ "in_index": true,
174
+ "tracks": [
175
+ {
176
+ "track": "RouterBench-0shot",
177
+ "score": 0.7864,
178
+ "headline": false
179
+ },
180
+ {
181
+ "track": "RouterBench-5shot",
182
+ "score": 0.8114,
183
+ "headline": false
184
+ }
185
+ ]
186
+ },
187
+ "9": {
188
+ "raw": 0.3063,
189
+ "skill": 0.3063,
190
+ "coverage": 1.0,
191
+ "random": 0.0,
192
+ "rule": "chance",
193
+ "in_index": true
194
+ },
195
+ "10": {
196
+ "raw": 0.1982,
197
+ "skill": 0.0,
198
+ "coverage": 1.0,
199
+ "random": 0.399,
200
+ "rule": "chance",
201
+ "in_index": true
202
+ },
203
+ "11": {
204
+ "raw": 0.8171,
205
+ "skill": 0.7355,
206
+ "coverage": 1.0,
207
+ "random": 0.3085,
208
+ "rule": "track",
209
+ "in_index": true,
210
+ "tracks": []
211
+ },
212
+ "12": {
213
+ "raw": 0.6899,
214
+ "skill": 0.5355,
215
+ "coverage": 1.0,
216
+ "random": 0.3324,
217
+ "rule": "chance",
218
+ "in_index": true
219
+ },
220
+ "20": {
221
+ "raw": 0.8408,
222
+ "skill": 0.6817,
223
+ "coverage": 1.0,
224
+ "random": 0.5,
225
+ "rule": "track",
226
+ "in_index": true,
227
+ "tracks": []
228
+ },
229
+ "21": {
230
+ "raw": 0.6381,
231
+ "skill": 0.2763,
232
+ "coverage": 1.0,
233
+ "random": 0.5,
234
+ "rule": "track",
235
+ "in_index": true,
236
+ "tracks": []
237
+ },
238
+ "22": {
239
+ "raw": 0.0763,
240
+ "skill": 0.0691,
241
+ "coverage": 1.0,
242
+ "random": 0.0078,
243
+ "rule": "track",
244
+ "in_index": true,
245
+ "tracks": []
246
+ },
247
+ "23": {
248
+ "raw": 0.5965,
249
+ "skill": 0.193,
250
+ "coverage": 1.0,
251
+ "random": 0.5,
252
+ "rule": "track",
253
+ "in_index": true,
254
+ "tracks": []
255
+ },
256
+ "25": {
257
+ "raw": 0.4082,
258
+ "skill": 0.2109,
259
+ "coverage": 1.0,
260
+ "random": 0.25,
261
+ "rule": "track",
262
+ "in_index": true,
263
+ "tracks": []
264
+ },
265
+ "28": {
266
+ "raw": 0.7419,
267
+ "skill": 0.4838,
268
+ "coverage": 1.0,
269
+ "random": 0.5,
270
+ "rule": "chance",
271
+ "in_index": true
272
+ },
273
+ "29": {
274
+ "raw": 0.867,
275
+ "skill": 0.8227,
276
+ "coverage": 1.0,
277
+ "random": 0.25,
278
+ "rule": "chance",
279
+ "in_index": true
280
+ },
281
+ "30": {
282
+ "raw": 0.6585,
283
+ "skill": 0.5882,
284
+ "coverage": 1.0,
285
+ "random": 0.25,
286
+ "rule": "track",
287
+ "in_index": true,
288
+ "tracks": [
289
+ {
290
+ "track": "GSM8K-4choice",
291
+ "score": 0.7096,
292
+ "headline": false
293
+ },
294
+ {
295
+ "track": "GSM8K-10choice",
296
+ "score": 0.6073,
297
+ "headline": false
298
+ }
299
+ ]
300
+ },
301
+ "31": {
302
+ "raw": 0.2292,
303
+ "skill": 0.1605,
304
+ "coverage": 1.0,
305
+ "random": 0.0819,
306
+ "rule": "track",
307
+ "in_index": true,
308
+ "tracks": []
309
+ },
310
+ "32": {
311
+ "raw": 0.5705,
312
+ "skill": 0.3172,
313
+ "coverage": 1.0,
314
+ "random": 0.371,
315
+ "rule": "chance",
316
+ "in_index": true
317
+ },
318
+ "33": {
319
+ "raw": 0.3467,
320
+ "skill": 0.338,
321
+ "coverage": 1.0,
322
+ "random": 0.0131,
323
+ "rule": "chance",
324
+ "in_index": true
325
+ },
326
+ "36": {
327
+ "raw": 0.1795,
328
+ "skill": 0.1396,
329
+ "coverage": 1.0,
330
+ "random": 0.0464,
331
+ "rule": "track",
332
+ "in_index": true,
333
+ "tracks": []
334
+ },
335
+ "37": {
336
+ "raw": 0.5199,
337
+ "skill": 0.3978,
338
+ "coverage": 1.0,
339
+ "random": 0.2027,
340
+ "rule": "track",
341
+ "in_index": true,
342
+ "tracks": []
343
+ },
344
+ "38": {
345
+ "raw": 0.055,
346
+ "skill": 0.055,
347
+ "coverage": 1.0,
348
+ "random": 0.0,
349
+ "rule": "chance",
350
+ "in_index": true
351
+ },
352
+ "39": {
353
+ "raw": 0.8246,
354
+ "skill": 0.742,
355
+ "coverage": 1.0,
356
+ "random": 0.3201,
357
+ "rule": "chance",
358
+ "in_index": true
359
+ },
360
+ "40": {
361
+ "raw": 0.5057,
362
+ "skill": 0.3641,
363
+ "coverage": 1.0,
364
+ "random": 0.2227,
365
+ "rule": "track",
366
+ "in_index": true,
367
+ "tracks": [
368
+ {
369
+ "track": "A \u00b7 Arabic",
370
+ "score": 0.3205,
371
+ "headline": false
372
+ },
373
+ {
374
+ "track": "A \u00b7 English",
375
+ "score": 0.5057,
376
+ "headline": true
377
+ },
378
+ {
379
+ "track": "C \u00b7 Arabic pairs",
380
+ "score": 0.79,
381
+ "headline": false
382
+ },
383
+ {
384
+ "track": "C \u00b7 English pairs",
385
+ "score": 0.95,
386
+ "headline": false
387
+ }
388
+ ]
389
+ },
390
+ "41": {
391
+ "raw": 0.7805,
392
+ "skill": 0.6707,
393
+ "coverage": 1.0,
394
+ "random": 0.3333,
395
+ "rule": "track",
396
+ "in_index": true,
397
+ "tracks": []
398
+ },
399
+ "42": {
400
+ "raw": 0.8038,
401
+ "skill": 0.6183,
402
+ "coverage": 1.0,
403
+ "random": 0.486,
404
+ "rule": "chance",
405
+ "in_index": true
406
+ },
407
+ "43": {
408
+ "raw": 0.5474,
409
+ "skill": 0.2819,
410
+ "coverage": 1.0,
411
+ "random": 0.3697,
412
+ "rule": "track",
413
+ "in_index": true,
414
+ "tracks": []
415
+ },
416
+ "44": {
417
+ "raw": 0.6614,
418
+ "skill": 0.3228,
419
+ "coverage": 1.0,
420
+ "random": 0.5,
421
+ "rule": "track",
422
+ "in_index": true,
423
+ "tracks": []
424
+ },
425
+ "45": {
426
+ "raw": 0.1098,
427
+ "skill": 0.0,
428
+ "coverage": 1.0,
429
+ "random": 0.1641,
430
+ "rule": "chance",
431
+ "in_index": true
432
+ },
433
+ "48": {
434
+ "raw": 0.282,
435
+ "skill": 0.282,
436
+ "coverage": 1.0,
437
+ "random": 0.25,
438
+ "rule": "vs baseline",
439
+ "in_index": true
440
+ },
441
+ "50": {
442
+ "raw": 0.4553,
443
+ "skill": 0.2094,
444
+ "coverage": 1.0,
445
+ "random": 0.311,
446
+ "rule": "track",
447
+ "in_index": true,
448
+ "tracks": []
449
+ },
450
+ "56": {
451
+ "raw": 0.6655,
452
+ "skill": 0.331,
453
+ "coverage": 1.0,
454
+ "random": 0.5,
455
+ "rule": "chance",
456
+ "in_index": true
457
+ },
458
+ "57": {
459
+ "raw": 0.6137,
460
+ "skill": 0.5655,
461
+ "coverage": 1.0,
462
+ "random": 0.1109,
463
+ "rule": "chance",
464
+ "in_index": true
465
+ },
466
+ "58": {
467
+ "raw": 0.6784,
468
+ "skill": 0.5338,
469
+ "coverage": 1.0,
470
+ "random": 0.3101,
471
+ "rule": "chance",
472
+ "in_index": true
473
+ },
474
+ "59": {
475
+ "raw": 0.6336,
476
+ "skill": 0.3776,
477
+ "coverage": 1.0,
478
+ "random": 0.4113,
479
+ "rule": "chance",
480
+ "in_index": true
481
+ },
482
+ "61": {
483
+ "raw": 0.6565,
484
+ "skill": 0.313,
485
+ "coverage": 1.0,
486
+ "random": 0.5,
487
+ "rule": "chance",
488
+ "in_index": true
489
+ },
490
+ "62": {
491
+ "raw": 0.6451,
492
+ "skill": 0.5268,
493
+ "coverage": 1.0,
494
+ "random": 0.25,
495
+ "rule": "chance",
496
+ "in_index": true
497
+ },
498
+ "64": {
499
+ "raw": 0.625,
500
+ "skill": 0.5312,
501
+ "coverage": 1.0,
502
+ "random": 0.2,
503
+ "rule": "chance",
504
+ "in_index": true
505
+ },
506
+ "24": {
507
+ "raw": 0.788,
508
+ "skill": 0.7173,
509
+ "coverage": 1.0,
510
+ "random": 0.25,
511
+ "rule": "shown, not counted",
512
+ "in_index": false
513
+ },
514
+ "26": {
515
+ "raw": 0.9798,
516
+ "skill": 0.9731,
517
+ "coverage": 1.0,
518
+ "random": 0.2502,
519
+ "rule": "shown, not counted",
520
+ "in_index": false
521
+ },
522
+ "27": {
523
+ "raw": 0.9514,
524
+ "skill": 0.9352,
525
+ "coverage": 1.0,
526
+ "random": 0.2502,
527
+ "rule": "shown, not counted",
528
+ "in_index": false
529
+ },
530
+ "34": {
531
+ "raw": 0.1,
532
+ "skill": 0.0,
533
+ "coverage": 1.0,
534
+ "random": 0.1667,
535
+ "rule": "shown, not counted",
536
+ "in_index": false
537
+ }
538
+ },
539
+ "benchmarks": {
540
+ "1": {
541
+ "catalog_id": 1,
542
+ "dataset": "BFCL",
543
+ "requests": 1694,
544
+ "answered": 1694,
545
+ "unsupported": 0,
546
+ "errors": 0,
547
+ "abstained": 0,
548
+ "pending": 0,
549
+ "scored_requests": 1694,
550
+ "metric": "case exact accuracy",
551
+ "score": 0.8926,
552
+ "reference_same_cases": null,
553
+ "median_ms": 139.8,
554
+ "index_raw": 0.8926,
555
+ "index_skill": 0.855,
556
+ "coverage": 1.0,
557
+ "chance": 0.2592,
558
+ "in_index": true
559
+ },
560
+ "2": {
561
+ "catalog_id": 2,
562
+ "dataset": "ToolRet",
563
+ "requests": 1000,
564
+ "answered": 1000,
565
+ "unsupported": 0,
566
+ "errors": 0,
567
+ "abstained": 0,
568
+ "pending": 0,
569
+ "scored_requests": 1000,
570
+ "metric": "nDCG@10",
571
+ "score": 0.4296,
572
+ "reference_same_cases": null,
573
+ "median_ms": 1187.9,
574
+ "index_raw": 0.4296,
575
+ "index_skill": 0.3719,
576
+ "coverage": 1.0,
577
+ "chance": 0.0918,
578
+ "in_index": true
579
+ },
580
+ "3": {
581
+ "catalog_id": 3,
582
+ "dataset": "API-Bank",
583
+ "requests": 508,
584
+ "answered": 508,
585
+ "unsupported": 0,
586
+ "errors": 0,
587
+ "abstained": 0,
588
+ "pending": 0,
589
+ "scored_requests": 508,
590
+ "metric": "accuracy",
591
+ "score": 0.8386,
592
+ "reference_same_cases": null,
593
+ "median_ms": 640.5,
594
+ "index_raw": 0.8386,
595
+ "index_skill": 0.8355,
596
+ "coverage": 1.0,
597
+ "chance": 0.0189,
598
+ "in_index": true
599
+ },
600
+ "4": {
601
+ "catalog_id": 4,
602
+ "dataset": "BANKING77",
603
+ "requests": 3080,
604
+ "answered": 3080,
605
+ "unsupported": 0,
606
+ "errors": 0,
607
+ "abstained": 0,
608
+ "pending": 0,
609
+ "scored_requests": 3080,
610
+ "metric": "macro-F1",
611
+ "score": 0.8177,
612
+ "reference_same_cases": null,
613
+ "median_ms": 103.1,
614
+ "index_raw": 0.8177,
615
+ "index_skill": 0.8154,
616
+ "coverage": 1.0,
617
+ "chance": 0.0127,
618
+ "in_index": true
619
+ },
620
+ "5": {
621
+ "catalog_id": 5,
622
+ "dataset": "CLINC150+OOS",
623
+ "requests": 5500,
624
+ "answered": 5500,
625
+ "unsupported": 0,
626
+ "errors": 0,
627
+ "abstained": 0,
628
+ "pending": 0,
629
+ "scored_requests": 5500,
630
+ "metric": "macro-F1",
631
+ "score": 0.8381,
632
+ "reference_same_cases": null,
633
+ "median_ms": 172.3,
634
+ "index_raw": 0.8381,
635
+ "index_skill": 0.8371,
636
+ "coverage": 1.0,
637
+ "chance": 0.006,
638
+ "in_index": true
639
+ },
640
+ "6": {
641
+ "catalog_id": 6,
642
+ "dataset": "RouterBench",
643
+ "requests": 10000,
644
+ "answered": 10000,
645
+ "unsupported": 0,
646
+ "errors": 0,
647
+ "abstained": 0,
648
+ "pending": 0,
649
+ "scored_requests": 10000,
650
+ "metric": "selected quality (quality objective)",
651
+ "score": 0.7989,
652
+ "reference_same_cases": null,
653
+ "median_ms": 392.5,
654
+ "tracks": [
655
+ {
656
+ "track": "RouterBench-0shot",
657
+ "score": 0.7864,
658
+ "headline": false
659
+ },
660
+ {
661
+ "track": "RouterBench-5shot",
662
+ "score": 0.8114,
663
+ "headline": false
664
+ }
665
+ ],
666
+ "index_raw": 0.7989,
667
+ "index_skill": 0.5256,
668
+ "coverage": 1.0,
669
+ "chance": 0.5245,
670
+ "in_index": true
671
+ },
672
+ "9": {
673
+ "catalog_id": 9,
674
+ "dataset": "Home appliance simulator",
675
+ "requests": 160,
676
+ "answered": 160,
677
+ "unsupported": 0,
678
+ "errors": 0,
679
+ "abstained": 0,
680
+ "pending": 0,
681
+ "scored_requests": 160,
682
+ "metric": "case exact accuracy",
683
+ "score": 0.3063,
684
+ "reference_same_cases": null,
685
+ "median_ms": 1713.5,
686
+ "index_raw": 0.3063,
687
+ "index_skill": 0.3063,
688
+ "coverage": 1.0,
689
+ "chance": 0.0,
690
+ "in_index": true
691
+ },
692
+ "10": {
693
+ "catalog_id": 10,
694
+ "dataset": "SGD/SGD-X",
695
+ "requests": 2500,
696
+ "answered": 2500,
697
+ "unsupported": 0,
698
+ "errors": 0,
699
+ "abstained": 0,
700
+ "pending": 0,
701
+ "scored_requests": 2500,
702
+ "metric": "macro-F1",
703
+ "score": 0.1982,
704
+ "reference_same_cases": null,
705
+ "median_ms": 106.5,
706
+ "index_raw": 0.1982,
707
+ "index_skill": 0.0,
708
+ "coverage": 1.0,
709
+ "chance": 0.399,
710
+ "in_index": true
711
+ },
712
+ "11": {
713
+ "catalog_id": 11,
714
+ "dataset": "ContractNLI",
715
+ "requests": 123,
716
+ "answered": 123,
717
+ "unsupported": 0,
718
+ "errors": 0,
719
+ "abstained": 0,
720
+ "pending": 0,
721
+ "scored_requests": 123,
722
+ "metric": "macro-F1",
723
+ "score": 0.8171,
724
+ "reference_same_cases": null,
725
+ "median_ms": 2966.1,
726
+ "index_raw": 0.8171,
727
+ "index_skill": 0.7355,
728
+ "coverage": 1.0,
729
+ "chance": 0.3085,
730
+ "in_index": true
731
+ },
732
+ "12": {
733
+ "catalog_id": 12,
734
+ "dataset": "ANLI",
735
+ "requests": 3200,
736
+ "answered": 3200,
737
+ "unsupported": 0,
738
+ "errors": 0,
739
+ "abstained": 0,
740
+ "pending": 0,
741
+ "scored_requests": 3200,
742
+ "metric": "macro-F1",
743
+ "score": 0.6899,
744
+ "reference_same_cases": null,
745
+ "median_ms": 19.4,
746
+ "index_raw": 0.6899,
747
+ "index_skill": 0.5355,
748
+ "coverage": 1.0,
749
+ "chance": 0.3324,
750
+ "in_index": true
751
+ },
752
+ "20": {
753
+ "catalog_id": 20,
754
+ "dataset": "BPoMP",
755
+ "requests": 5000,
756
+ "answered": 5000,
757
+ "unsupported": 0,
758
+ "errors": 0,
759
+ "abstained": 0,
760
+ "pending": 0,
761
+ "scored_requests": 5000,
762
+ "metric": "accuracy",
763
+ "score": 0.8398,
764
+ "reference_same_cases": null,
765
+ "median_ms": 18.7,
766
+ "index_raw": 0.8408,
767
+ "index_skill": 0.6817,
768
+ "coverage": 1.0,
769
+ "chance": 0.5,
770
+ "in_index": true
771
+ },
772
+ "21": {
773
+ "catalog_id": 21,
774
+ "dataset": "Humicroedit",
775
+ "requests": 2628,
776
+ "answered": 2628,
777
+ "unsupported": 0,
778
+ "errors": 0,
779
+ "abstained": 0,
780
+ "pending": 0,
781
+ "scored_requests": 2628,
782
+ "metric": "accuracy",
783
+ "score": 0.6381,
784
+ "reference_same_cases": null,
785
+ "median_ms": 11.8,
786
+ "index_raw": 0.6381,
787
+ "index_skill": 0.2763,
788
+ "coverage": 1.0,
789
+ "chance": 0.5,
790
+ "in_index": true
791
+ },
792
+ "22": {
793
+ "catalog_id": 22,
794
+ "dataset": "POP909-CL",
795
+ "requests": 2000,
796
+ "answered": 2000,
797
+ "unsupported": 0,
798
+ "errors": 0,
799
+ "abstained": 0,
800
+ "pending": 0,
801
+ "scored_requests": 2000,
802
+ "metric": "accuracy",
803
+ "score": 0.0825,
804
+ "reference_same_cases": null,
805
+ "median_ms": 843.5,
806
+ "index_raw": 0.0763,
807
+ "index_skill": 0.0691,
808
+ "coverage": 1.0,
809
+ "chance": 0.0078,
810
+ "in_index": true
811
+ },
812
+ "23": {
813
+ "catalog_id": 23,
814
+ "dataset": "cfcolor",
815
+ "requests": 5000,
816
+ "answered": 5000,
817
+ "unsupported": 0,
818
+ "errors": 0,
819
+ "abstained": 0,
820
+ "pending": 0,
821
+ "scored_requests": 5000,
822
+ "metric": "accuracy",
823
+ "score": 0.5886,
824
+ "reference_same_cases": null,
825
+ "median_ms": 50.0,
826
+ "index_raw": 0.5965,
827
+ "index_skill": 0.193,
828
+ "coverage": 1.0,
829
+ "chance": 0.5,
830
+ "in_index": true
831
+ },
832
+ "24": {
833
+ "catalog_id": 24,
834
+ "dataset": "MMLU",
835
+ "requests": 14033,
836
+ "answered": 14033,
837
+ "unsupported": 0,
838
+ "errors": 0,
839
+ "abstained": 0,
840
+ "pending": 0,
841
+ "scored_requests": 14033,
842
+ "metric": "accuracy",
843
+ "score": 0.788,
844
+ "reference_same_cases": null,
845
+ "median_ms": 18.2,
846
+ "index_raw": 0.788,
847
+ "index_skill": 0.7173,
848
+ "coverage": 1.0,
849
+ "chance": 0.25,
850
+ "in_index": false
851
+ },
852
+ "25": {
853
+ "catalog_id": 25,
854
+ "dataset": "GPQA Diamond",
855
+ "requests": 196,
856
+ "answered": 196,
857
+ "unsupported": 0,
858
+ "errors": 0,
859
+ "abstained": 0,
860
+ "pending": 0,
861
+ "scored_requests": 196,
862
+ "metric": "accuracy",
863
+ "score": 0.4082,
864
+ "reference_same_cases": null,
865
+ "median_ms": 34.6,
866
+ "index_raw": 0.4082,
867
+ "index_skill": 0.2109,
868
+ "coverage": 1.0,
869
+ "chance": 0.25,
870
+ "in_index": true
871
+ },
872
+ "26": {
873
+ "catalog_id": 26,
874
+ "dataset": "ARC-Easy",
875
+ "requests": 2376,
876
+ "answered": 2376,
877
+ "unsupported": 0,
878
+ "errors": 0,
879
+ "abstained": 0,
880
+ "pending": 0,
881
+ "scored_requests": 2376,
882
+ "metric": "accuracy",
883
+ "score": 0.9798,
884
+ "reference_same_cases": null,
885
+ "median_ms": 16.1,
886
+ "index_raw": 0.9798,
887
+ "index_skill": 0.9731,
888
+ "coverage": 1.0,
889
+ "chance": 0.2502,
890
+ "in_index": false
891
+ },
892
+ "27": {
893
+ "catalog_id": 27,
894
+ "dataset": "ARC-Challenge",
895
+ "requests": 1172,
896
+ "answered": 1172,
897
+ "unsupported": 0,
898
+ "errors": 0,
899
+ "abstained": 0,
900
+ "pending": 0,
901
+ "scored_requests": 1172,
902
+ "metric": "accuracy",
903
+ "score": 0.9514,
904
+ "reference_same_cases": null,
905
+ "median_ms": 16.9,
906
+ "index_raw": 0.9514,
907
+ "index_skill": 0.9352,
908
+ "coverage": 1.0,
909
+ "chance": 0.2502,
910
+ "in_index": false
911
+ },
912
+ "28": {
913
+ "catalog_id": 28,
914
+ "dataset": "WinoGrande",
915
+ "requests": 1267,
916
+ "answered": 1267,
917
+ "unsupported": 0,
918
+ "errors": 0,
919
+ "abstained": 0,
920
+ "pending": 0,
921
+ "scored_requests": 1267,
922
+ "metric": "accuracy",
923
+ "score": 0.7419,
924
+ "reference_same_cases": null,
925
+ "median_ms": 12.0,
926
+ "index_raw": 0.7419,
927
+ "index_skill": 0.4838,
928
+ "coverage": 1.0,
929
+ "chance": 0.5,
930
+ "in_index": true
931
+ },
932
+ "29": {
933
+ "catalog_id": 29,
934
+ "dataset": "HellaSwag",
935
+ "requests": 10042,
936
+ "answered": 10042,
937
+ "unsupported": 0,
938
+ "errors": 0,
939
+ "abstained": 0,
940
+ "pending": 0,
941
+ "scored_requests": 10042,
942
+ "metric": "accuracy",
943
+ "score": 0.867,
944
+ "reference_same_cases": null,
945
+ "median_ms": 27.6,
946
+ "index_raw": 0.867,
947
+ "index_skill": 0.8227,
948
+ "coverage": 1.0,
949
+ "chance": 0.25,
950
+ "in_index": true
951
+ },
952
+ "30": {
953
+ "catalog_id": 30,
954
+ "dataset": "GSM8K",
955
+ "requests": 2638,
956
+ "answered": 2638,
957
+ "unsupported": 0,
958
+ "errors": 0,
959
+ "abstained": 0,
960
+ "pending": 0,
961
+ "scored_requests": 2638,
962
+ "metric": "accuracy",
963
+ "score": 0.6585,
964
+ "reference_same_cases": null,
965
+ "median_ms": 24.3,
966
+ "tracks": [
967
+ {
968
+ "track": "GSM8K-4choice",
969
+ "score": 0.7096,
970
+ "headline": false
971
+ },
972
+ {
973
+ "track": "GSM8K-10choice",
974
+ "score": 0.6073,
975
+ "headline": false
976
+ }
977
+ ],
978
+ "index_raw": 0.6585,
979
+ "index_skill": 0.5882,
980
+ "coverage": 1.0,
981
+ "chance": 0.25,
982
+ "in_index": true
983
+ },
984
+ "31": {
985
+ "catalog_id": 31,
986
+ "dataset": "ChessBench",
987
+ "requests": 5000,
988
+ "answered": 5000,
989
+ "unsupported": 0,
990
+ "errors": 0,
991
+ "abstained": 0,
992
+ "pending": 0,
993
+ "scored_requests": 5000,
994
+ "metric": "accuracy",
995
+ "score": 0.2292,
996
+ "reference_same_cases": null,
997
+ "median_ms": 132.4,
998
+ "index_raw": 0.2292,
999
+ "index_skill": 0.1605,
1000
+ "coverage": 1.0,
1001
+ "chance": 0.0819,
1002
+ "in_index": true
1003
+ },
1004
+ "32": {
1005
+ "catalog_id": 32,
1006
+ "dataset": "MuSR",
1007
+ "requests": 752,
1008
+ "answered": 752,
1009
+ "unsupported": 0,
1010
+ "errors": 0,
1011
+ "abstained": 0,
1012
+ "pending": 0,
1013
+ "scored_requests": 752,
1014
+ "metric": "accuracy",
1015
+ "score": 0.5705,
1016
+ "reference_same_cases": null,
1017
+ "median_ms": 93.7,
1018
+ "index_raw": 0.5705,
1019
+ "index_skill": 0.3172,
1020
+ "coverage": 1.0,
1021
+ "chance": 0.371,
1022
+ "in_index": true
1023
+ },
1024
+ "33": {
1025
+ "catalog_id": 33,
1026
+ "dataset": "SATA-Bench",
1027
+ "requests": 1650,
1028
+ "answered": 1650,
1029
+ "unsupported": 0,
1030
+ "errors": 0,
1031
+ "abstained": 0,
1032
+ "pending": 0,
1033
+ "scored_requests": 1650,
1034
+ "metric": "case exact accuracy",
1035
+ "score": 0.3467,
1036
+ "reference_same_cases": null,
1037
+ "median_ms": 321.2,
1038
+ "index_raw": 0.3467,
1039
+ "index_skill": 0.338,
1040
+ "coverage": 1.0,
1041
+ "chance": 0.0131,
1042
+ "in_index": true
1043
+ },
1044
+ "34": {
1045
+ "catalog_id": 34,
1046
+ "dataset": "SimpleBench",
1047
+ "requests": 10,
1048
+ "answered": 10,
1049
+ "unsupported": 0,
1050
+ "errors": 0,
1051
+ "abstained": 0,
1052
+ "pending": 0,
1053
+ "scored_requests": 10,
1054
+ "metric": "accuracy",
1055
+ "score": 0.1,
1056
+ "reference_same_cases": null,
1057
+ "median_ms": 21.2,
1058
+ "index_raw": 0.1,
1059
+ "index_skill": 0.0,
1060
+ "coverage": 1.0,
1061
+ "chance": 0.1667,
1062
+ "in_index": false
1063
+ },
1064
+ "36": {
1065
+ "catalog_id": 36,
1066
+ "dataset": "BRIGHT",
1067
+ "requests": 550,
1068
+ "answered": 550,
1069
+ "unsupported": 0,
1070
+ "errors": 0,
1071
+ "abstained": 0,
1072
+ "pending": 0,
1073
+ "scored_requests": 550,
1074
+ "metric": "nDCG@10",
1075
+ "score": 0.1795,
1076
+ "reference_same_cases": null,
1077
+ "median_ms": 1508.7,
1078
+ "index_raw": 0.1795,
1079
+ "index_skill": 0.1396,
1080
+ "coverage": 1.0,
1081
+ "chance": 0.0464,
1082
+ "in_index": true
1083
+ },
1084
+ "37": {
1085
+ "catalog_id": 37,
1086
+ "dataset": "Amazon ESCI",
1087
+ "requests": 5000,
1088
+ "answered": 5000,
1089
+ "unsupported": 0,
1090
+ "errors": 0,
1091
+ "abstained": 0,
1092
+ "pending": 0,
1093
+ "scored_requests": 5000,
1094
+ "metric": "macro-F1",
1095
+ "score": 0.5199,
1096
+ "reference_same_cases": null,
1097
+ "median_ms": 39.0,
1098
+ "index_raw": 0.5199,
1099
+ "index_skill": 0.3978,
1100
+ "coverage": 1.0,
1101
+ "chance": 0.2027,
1102
+ "in_index": true
1103
+ },
1104
+ "38": {
1105
+ "catalog_id": 38,
1106
+ "dataset": "ACOS",
1107
+ "requests": 1565,
1108
+ "answered": 1565,
1109
+ "unsupported": 0,
1110
+ "errors": 0,
1111
+ "abstained": 0,
1112
+ "pending": 0,
1113
+ "scored_requests": 1565,
1114
+ "metric": "case exact accuracy",
1115
+ "score": 0.055,
1116
+ "reference_same_cases": null,
1117
+ "median_ms": 982.6,
1118
+ "index_raw": 0.055,
1119
+ "index_skill": 0.055,
1120
+ "coverage": 1.0,
1121
+ "chance": 0.0,
1122
+ "in_index": true
1123
+ },
1124
+ "39": {
1125
+ "catalog_id": 39,
1126
+ "dataset": "FinEntity",
1127
+ "requests": 979,
1128
+ "answered": 979,
1129
+ "unsupported": 0,
1130
+ "errors": 0,
1131
+ "abstained": 0,
1132
+ "pending": 0,
1133
+ "scored_requests": 979,
1134
+ "metric": "macro-F1",
1135
+ "score": 0.8246,
1136
+ "reference_same_cases": null,
1137
+ "median_ms": 36.0,
1138
+ "index_raw": 0.8246,
1139
+ "index_skill": 0.742,
1140
+ "coverage": 1.0,
1141
+ "chance": 0.3201,
1142
+ "in_index": true
1143
+ },
1144
+ "40": {
1145
+ "catalog_id": 40,
1146
+ "dataset": "iSarcasmEval",
1147
+ "requests": 4600,
1148
+ "answered": 4600,
1149
+ "unsupported": 0,
1150
+ "errors": 0,
1151
+ "abstained": 0,
1152
+ "pending": 0,
1153
+ "scored_requests": 4600,
1154
+ "metric": "Sarcasm F1 \u00b7 track A, English",
1155
+ "score": 0.5057,
1156
+ "reference_same_cases": null,
1157
+ "median_ms": 12.9,
1158
+ "tracks": [
1159
+ {
1160
+ "track": "A \u00b7 Arabic",
1161
+ "score": 0.3205,
1162
+ "headline": false
1163
+ },
1164
+ {
1165
+ "track": "A \u00b7 English",
1166
+ "score": 0.5057,
1167
+ "headline": true
1168
+ },
1169
+ {
1170
+ "track": "C \u00b7 Arabic pairs",
1171
+ "score": 0.79,
1172
+ "headline": false
1173
+ },
1174
+ {
1175
+ "track": "C \u00b7 English pairs",
1176
+ "score": 0.95,
1177
+ "headline": false
1178
+ }
1179
+ ],
1180
+ "index_raw": 0.5057,
1181
+ "index_skill": 0.3641,
1182
+ "coverage": 1.0,
1183
+ "chance": 0.2227,
1184
+ "in_index": true
1185
+ },
1186
+ "41": {
1187
+ "catalog_id": 41,
1188
+ "dataset": "VAST",
1189
+ "requests": 3006,
1190
+ "answered": 3006,
1191
+ "unsupported": 0,
1192
+ "errors": 0,
1193
+ "abstained": 0,
1194
+ "pending": 0,
1195
+ "scored_requests": 3006,
1196
+ "metric": "macro-F1",
1197
+ "score": 0.7805,
1198
+ "reference_same_cases": null,
1199
+ "median_ms": 23.6,
1200
+ "index_raw": 0.7805,
1201
+ "index_skill": 0.6707,
1202
+ "coverage": 1.0,
1203
+ "chance": 0.3333,
1204
+ "in_index": true
1205
+ },
1206
+ "42": {
1207
+ "catalog_id": 42,
1208
+ "dataset": "NLI4CT",
1209
+ "requests": 5500,
1210
+ "answered": 5500,
1211
+ "unsupported": 0,
1212
+ "errors": 0,
1213
+ "abstained": 0,
1214
+ "pending": 0,
1215
+ "scored_requests": 5500,
1216
+ "metric": "macro-F1",
1217
+ "score": 0.8038,
1218
+ "reference_same_cases": null,
1219
+ "median_ms": 54.9,
1220
+ "index_raw": 0.8038,
1221
+ "index_skill": 0.6183,
1222
+ "coverage": 1.0,
1223
+ "chance": 0.486,
1224
+ "in_index": true
1225
+ },
1226
+ "43": {
1227
+ "catalog_id": 43,
1228
+ "dataset": "CRUXEval",
1229
+ "requests": 570,
1230
+ "answered": 570,
1231
+ "unsupported": 0,
1232
+ "errors": 0,
1233
+ "abstained": 0,
1234
+ "pending": 0,
1235
+ "scored_requests": 570,
1236
+ "metric": "accuracy",
1237
+ "score": 0.5474,
1238
+ "reference_same_cases": null,
1239
+ "median_ms": 20.4,
1240
+ "index_raw": 0.5474,
1241
+ "index_skill": 0.2819,
1242
+ "coverage": 1.0,
1243
+ "chance": 0.3697,
1244
+ "in_index": true
1245
+ },
1246
+ "44": {
1247
+ "catalog_id": 44,
1248
+ "dataset": "CLadder",
1249
+ "requests": 5000,
1250
+ "answered": 5000,
1251
+ "unsupported": 0,
1252
+ "errors": 0,
1253
+ "abstained": 0,
1254
+ "pending": 0,
1255
+ "scored_requests": 5000,
1256
+ "metric": "accuracy",
1257
+ "score": 0.6614,
1258
+ "reference_same_cases": null,
1259
+ "median_ms": 19.9,
1260
+ "index_raw": 0.6614,
1261
+ "index_skill": 0.3228,
1262
+ "coverage": 1.0,
1263
+ "chance": 0.5,
1264
+ "in_index": true
1265
+ },
1266
+ "45": {
1267
+ "catalog_id": 45,
1268
+ "dataset": "HLE",
1269
+ "requests": 501,
1270
+ "answered": 501,
1271
+ "unsupported": 0,
1272
+ "errors": 0,
1273
+ "abstained": 0,
1274
+ "pending": 0,
1275
+ "scored_requests": 501,
1276
+ "metric": "accuracy",
1277
+ "score": 0.1098,
1278
+ "reference_same_cases": null,
1279
+ "median_ms": 30.4,
1280
+ "index_raw": 0.1098,
1281
+ "index_skill": 0.0,
1282
+ "coverage": 1.0,
1283
+ "chance": 0.1641,
1284
+ "in_index": true
1285
+ },
1286
+ "48": {
1287
+ "catalog_id": 48,
1288
+ "dataset": "ForecastBench",
1289
+ "requests": 10139,
1290
+ "answered": 10139,
1291
+ "unsupported": 0,
1292
+ "errors": 0,
1293
+ "abstained": 0,
1294
+ "pending": 0,
1295
+ "scored_requests": 10139,
1296
+ "metric": "Brier (lower is better)",
1297
+ "score": 0.1795,
1298
+ "reference_same_cases": null,
1299
+ "median_ms": 73.2,
1300
+ "index_raw": 0.282,
1301
+ "index_skill": 0.282,
1302
+ "coverage": 1.0,
1303
+ "chance": 0.25,
1304
+ "in_index": true
1305
+ },
1306
+ "50": {
1307
+ "catalog_id": 50,
1308
+ "dataset": "Habermas Machine",
1309
+ "requests": 1676,
1310
+ "answered": 1676,
1311
+ "unsupported": 0,
1312
+ "errors": 0,
1313
+ "abstained": 0,
1314
+ "pending": 0,
1315
+ "scored_requests": 1676,
1316
+ "metric": "accuracy",
1317
+ "score": 0.4553,
1318
+ "reference_same_cases": null,
1319
+ "median_ms": 81.4,
1320
+ "index_raw": 0.4553,
1321
+ "index_skill": 0.2094,
1322
+ "coverage": 1.0,
1323
+ "chance": 0.311,
1324
+ "in_index": true
1325
+ },
1326
+ "56": {
1327
+ "catalog_id": 56,
1328
+ "dataset": "PhishNChips phishing decisions",
1329
+ "requests": 2000,
1330
+ "answered": 2000,
1331
+ "unsupported": 0,
1332
+ "errors": 0,
1333
+ "abstained": 0,
1334
+ "pending": 0,
1335
+ "metric": "accuracy",
1336
+ "score": 0.6655,
1337
+ "median_ms": 208.8,
1338
+ "scored_requests": 2000,
1339
+ "index_raw": 0.6655,
1340
+ "index_skill": 0.331,
1341
+ "coverage": 1.0,
1342
+ "chance": 0.5,
1343
+ "in_index": true
1344
+ },
1345
+ "57": {
1346
+ "catalog_id": 57,
1347
+ "dataset": "MMLU-Pro",
1348
+ "requests": 12032,
1349
+ "answered": 12032,
1350
+ "unsupported": 0,
1351
+ "errors": 0,
1352
+ "abstained": 0,
1353
+ "pending": 0,
1354
+ "metric": "accuracy",
1355
+ "score": 0.6137,
1356
+ "median_ms": 30.6,
1357
+ "scored_requests": 12032,
1358
+ "index_raw": 0.6137,
1359
+ "index_skill": 0.5655,
1360
+ "coverage": 1.0,
1361
+ "chance": 0.1109,
1362
+ "in_index": true
1363
+ },
1364
+ "58": {
1365
+ "catalog_id": 58,
1366
+ "dataset": "BBH fixed-option tasks",
1367
+ "requests": 5507,
1368
+ "answered": 5507,
1369
+ "unsupported": 0,
1370
+ "errors": 0,
1371
+ "abstained": 0,
1372
+ "pending": 0,
1373
+ "metric": "accuracy",
1374
+ "score": 0.6784,
1375
+ "median_ms": 23.7,
1376
+ "scored_requests": 5507,
1377
+ "index_raw": 0.6784,
1378
+ "index_skill": 0.5338,
1379
+ "coverage": 1.0,
1380
+ "chance": 0.3101,
1381
+ "in_index": true
1382
+ },
1383
+ "59": {
1384
+ "catalog_id": 59,
1385
+ "dataset": "RAGTruth response-level hallucination",
1386
+ "requests": 2700,
1387
+ "answered": 2700,
1388
+ "unsupported": 0,
1389
+ "errors": 0,
1390
+ "abstained": 0,
1391
+ "pending": 0,
1392
+ "metric": "F1 on hallucinated class",
1393
+ "score": 0.6336,
1394
+ "median_ms": 77.1,
1395
+ "scored_requests": 2700,
1396
+ "index_raw": 0.6336,
1397
+ "index_skill": 0.3776,
1398
+ "coverage": 1.0,
1399
+ "chance": 0.4113,
1400
+ "in_index": true
1401
+ },
1402
+ "61": {
1403
+ "catalog_id": 61,
1404
+ "dataset": "HoVer claim verification",
1405
+ "requests": 4000,
1406
+ "answered": 4000,
1407
+ "unsupported": 0,
1408
+ "errors": 0,
1409
+ "abstained": 0,
1410
+ "pending": 0,
1411
+ "metric": "accuracy",
1412
+ "score": 0.6565,
1413
+ "median_ms": 45.6,
1414
+ "scored_requests": 4000,
1415
+ "index_raw": 0.6565,
1416
+ "index_skill": 0.313,
1417
+ "coverage": 1.0,
1418
+ "chance": 0.5,
1419
+ "in_index": true
1420
+ },
1421
+ "62": {
1422
+ "catalog_id": 62,
1423
+ "dataset": "When2Call MCQ",
1424
+ "requests": 3652,
1425
+ "answered": 3652,
1426
+ "unsupported": 0,
1427
+ "errors": 0,
1428
+ "abstained": 0,
1429
+ "pending": 0,
1430
+ "metric": "accuracy",
1431
+ "score": 0.6451,
1432
+ "median_ms": 79.8,
1433
+ "scored_requests": 3652,
1434
+ "index_raw": 0.6451,
1435
+ "index_skill": 0.5268,
1436
+ "coverage": 1.0,
1437
+ "chance": 0.25,
1438
+ "in_index": true
1439
+ },
1440
+ "64": {
1441
+ "catalog_id": 64,
1442
+ "dataset": "New Yorker caption matching",
1443
+ "requests": 528,
1444
+ "answered": 528,
1445
+ "unsupported": 0,
1446
+ "errors": 0,
1447
+ "abstained": 0,
1448
+ "pending": 0,
1449
+ "metric": "accuracy",
1450
+ "score": 0.625,
1451
+ "median_ms": 25.2,
1452
+ "scored_requests": 528,
1453
+ "index_raw": 0.625,
1454
+ "index_skill": 0.5312,
1455
+ "coverage": 1.0,
1456
+ "chance": 0.2,
1457
+ "in_index": true
1458
+ }
1459
+ },
1460
+ "panel_id": "decision-index-0.2",
1461
+ "note": "Decision Index 0.2 averages 40 benchmarks in five equal-weight areas: each area is the plain mean of its benchmarks and the index is 100 x the mean of the five areas. Each benchmark is chance-corrected first, (score - chance) / (1 - chance) clipped to 0-1, so 0 means random guessing and 100 means perfect. Every score is coverage-adjusted, so an unanswered or unsupported request counts as wrong. ForecastBench enters against its baseline: clip((0.25 - Brier) / 0.25) x coverage, so always predicting 0.5 scores zero. MMLU, ARC-Easy, ARC-Challenge, SimpleBench stay on the board as non-index benchmarks. The six interactive environments are still unrun and stay out. Every entrant on the board has results on all 40 index benchmarks. Point estimates only, no uncertainty intervals yet.",
1462
+ "local_run": {
1463
+ "kit_commit": "19ad28ec9485493cc4f7fc07d91c178f948e6434",
1464
+ "evaluated": "2026-09-25",
1465
+ "note": "Scored locally with the official kit; not a leaderboard submission or result. Requests shared with 0.1 reuse this model's 0.1 predictions; the added requests ran with the same frozen evaluation setup at temperature 1.0. No leaderboard-style exposure penalty is applied: known training exposure (see the model card) stays in these scores.",
1466
+ "without_mmlu_pro": 42.84,
1467
+ "screening": [
1468
+ {
1469
+ "stage": "MiMo",
1470
+ "rows": 123195,
1471
+ "matched": 180,
1472
+ "by_source": {
1473
+ "MMLU-Pro": 0,
1474
+ "SuperGPQA": 176,
1475
+ "BoolQ": 3,
1476
+ "MedMCQA": 1
1477
+ }
1478
+ }
1479
+ ],
1480
+ "timing_note": "latency_ms and every benchmark's median_ms are each request's share of batched inference time, allocated by prompt tokens. They are not serial-request or HTTP-serving latency and shouldn't be used for serving-latency comparisons."
1481
+ }
1482
+ }