Image-Text-to-Text
Transformers
Safetensors
English
qwen3_5
decision-model
typed-decisions
one-pass
option-probabilities
conversational
Instructions to use thegovind/blink-mimo-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use thegovind/blink-mimo-9b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="thegovind/blink-mimo-9b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("thegovind/blink-mimo-9b") model = AutoModelForMultimodalLM.from_pretrained("thegovind/blink-mimo-9b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use thegovind/blink-mimo-9b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "thegovind/blink-mimo-9b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thegovind/blink-mimo-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/thegovind/blink-mimo-9b
- SGLang
How to use thegovind/blink-mimo-9b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "thegovind/blink-mimo-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thegovind/blink-mimo-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "thegovind/blink-mimo-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thegovind/blink-mimo-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use thegovind/blink-mimo-9b with Docker Model Runner:
docker model run hf.co/thegovind/blink-mimo-9b
Card: local Decision Index 0.2 run (descriptive; known training exposure not penalized) and its evidence file
Browse files- README.md +51 -4
- eval/decision-index-0.2-local.json +1482 -0
README.md
CHANGED
|
@@ -32,6 +32,55 @@ Xiaomi, Alibaba Cloud or the Qwen team. Weights are for non-commercial research;
|
|
| 32 |
|
| 33 |
## Results
|
| 34 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
### Decision Index 0.1 (archived edition)
|
| 36 |
|
| 37 |
**Full suite: blink-mimo-9b 56.53 vs Jev 1.13.0 59.51.**
|
|
@@ -48,9 +97,7 @@ Xiaomi, Alibaba Cloud or the Qwen team. Weights are for non-commercial research;
|
|
| 48 |
| Kev 9B | 9B | 50.48 | 32.96 | 30.54 |
|
| 49 |
| Kev 4B | 4B | 47.43 | 28.86 | 25.67 |
|
| 50 |
|
| 51 |
-
We ran the complete archived 0.1 suite: 132,422 requests across 37 benchmarks. The headline index averages 19 panel benchmarks. Comparison rows use the 2026-09-22 leaderboard snapshot. We ran the official kit's scorer locally; these aren't leaderboard submissions. The live [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index) moved to 0.2 on 2026-09-24
|
| 52 |
-
|
| 53 |
-
No Decision Index 0.2 result is reported for these models. Comparable shared-benchmark results require matched request subsets and the 0.2 metric transformations.
|
| 54 |
|
| 55 |
Third in our local archived 0.1 comparison, behind blink-27b and Jev and ahead of every open entry in the September 22 snapshot (best: Jevfire).
|
| 56 |
|
|
@@ -287,7 +334,7 @@ These are source-repository licences; they don't settle rights in every underlyi
|
|
| 287 |
- **Training overlap.** Public train splits also used by the 0.1 index: ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, ChessBench (searchless_chess training positions; none of the 5,000 test positions), GSM8K (train split; solution-checking items). We also used ANLI and BANKING77 train splits; they're in the 0.1 suite but outside its index, and both are in the 0.2 panel. The audit below reports what was checked and any shared passages.
|
| 288 |
- **Partitions.** Public-source data included training and development partitions.
|
| 289 |
- **Final-mixture audit.** Rechecked every question row (including teacher-written rows) against the complete 0.1 suite (132,422 requests) and JevBench's 231 public items. The checks looked for exact matches of normalised strings of at least 30 characters in any field and shared 13-word passages in each row's question text (instructions, state.question, state.code). Strings or passages seen in 20 or more suite requests were treated as prompt templates and ignored. No public JevBench item matched under these checks; a separate position check found no shared chess positions.
|
| 290 |
-
- **Suite overlap.** 16 BANKING77/VAST training rows share a 13-word passage with 31 suite requests: 23 of VAST's 3,006 and 8 of BANKING77's 3,080. Two VAST training posts are near-duplicates of a test post, but these rows had no exact normalised-text match under the audit. Dropping those requests leaves the index at 56.53 (VAST 0.7805 → 0.7803); BANKING77 is outside the index.
|
| 291 |
- **Audit limits.** The 13-word passage check didn't search long-document bodies or option text. Semantic or pretraining overlap can't be ruled out, and private JevBench items weren't available to check.
|
| 292 |
- **Generated reasoning.** Our programs computed the labels for CRUXEval-style code and CLadder-style causal questions; no items from those benchmarks were used. We didn't reuse the suite's GSM8K distractors.
|
| 293 |
- **Teacher documents.** We kept Qwen3.8-27B's documents only if a fresh blind solve by that same teacher agreed with the answer. That's an agreement filter, not independent verification.
|
|
|
|
| 32 |
|
| 33 |
## Results
|
| 34 |
|
| 35 |
+
### Decision Index 0.2 (local run)
|
| 36 |
+
|
| 37 |
+
We ran the full Decision Index 0.2 suite ourselves with the official scoring kit at commit 19ad28e on 2026-09-25. This is a descriptive run, not a leaderboard submission or accepted result. The kit's scorer does not apply the leaderboard's penalty for rows an entrant trained on, so known training exposure stays in these scores and they cannot be ranked against the leaderboard.
|
| 38 |
+
|
| 39 |
+
| Balanced skill | Balanced raw | Breadth skill | Without MMLU-Pro |
|
| 40 |
+
|---:|---:|---:|---:|
|
| 41 |
+
| 43.36 | 57.24 | 42.38 | 42.84 |
|
| 42 |
+
|
| 43 |
+
- Our Blink fine-tune did not use MMLU-Pro or GPQA as direct data sources. Screening found 176 SuperGPQA training rows matching added-request text. SuperGPQA is a training source, not a benchmark in this suite. Public train splits used in training are listed under “Training overlap” below.
|
| 44 |
+
|
| 45 |
+
<details><summary>Extra tables and method</summary>
|
| 46 |
+
|
| 47 |
+
| Area | Number of benchmarks | Skill | Raw |
|
| 48 |
+
|---|---:|---:|---:|
|
| 49 |
+
| Knowledge & Reasoning | 10 | 33.2 | 48.2 |
|
| 50 |
+
| Language Understanding | 10 | 54.1 | 67.2 |
|
| 51 |
+
| Retrieval & Classification | 7 | 40.5 | 55.4 |
|
| 52 |
+
| Tools & Automation | 6 | 57.0 | 65.2 |
|
| 53 |
+
| Arts & Human Taste | 7 | 32.0 | 50.2 |
|
| 54 |
+
|
| 55 |
+
The seven benchmarks added in 0.2.
|
| 56 |
+
|
| 57 |
+
| Benchmark | Metric | Requests | Answered | Raw | Skill |
|
| 58 |
+
|---|---|---:|---:|---:|---:|
|
| 59 |
+
| PhishNChips phishing decisions | accuracy | 2,000 | 2,000 | 66.5 | 33.1 |
|
| 60 |
+
| MMLU-Pro | accuracy | 12,032 | 12,032 | 61.4 | 56.5 |
|
| 61 |
+
| BBH fixed-option tasks | accuracy | 5,507 | 5,507 | 67.8 | 53.4 |
|
| 62 |
+
| RAGTruth response-level hallucination | F1 on hallucinated class | 2,700 | 2,700 | 63.4 | 37.8 |
|
| 63 |
+
| HoVer claim verification | accuracy | 4,000 | 4,000 | 65.6 | 31.3 |
|
| 64 |
+
| When2Call MCQ | accuracy | 3,652 | 3,652 | 64.5 | 52.7 |
|
| 65 |
+
| New Yorker caption matching | accuracy | 528 | 528 | 62.5 | 53.1 |
|
| 66 |
+
|
| 67 |
+
- All 151,034 of 151,034 scoreable requests scored. Of 44 scored benchmarks, 40 count toward the index across five equal areas.
|
| 68 |
+
- Requests shared with 0.1 reuse the model's 0.1 predictions. We ran the 30,419 added requests with the same frozen evaluation setup as 0.1, the evaluated adapter loaded on the MiMo base, which the published graft matched on JevBench's 231 public items (see Evaluation notes), at temperature 1.0.
|
| 69 |
+
- Balanced skill is the headline index. “Without MMLU-Pro” drops MMLU-Pro, averages the other nine Knowledge benchmarks, and keeps five equal areas. It is a sensitivity check, not a score free of training effects.
|
| 70 |
+
- These are point estimates, with no significance, calibration, or latency claims. Do not compare them with 0.1 numbers because the editions differ.
|
| 71 |
+
|
| 72 |
+
Training-row text matches in the added requests.
|
| 73 |
+
|
| 74 |
+
| Training stage | Rows in the stage | Rows matching added-request text | From MMLU-Pro | From SuperGPQA | Other |
|
| 75 |
+
|---|---:|---:|---:|---:|---:|
|
| 76 |
+
| MiMo | 123,195 | 180 | 0 | 176 | 4 |
|
| 77 |
+
|
| 78 |
+
We screened for exact normalised strings of at least 30 characters shared by training rows and added requests, ignoring strings found in 20 or more requests as templates. Counts are training rows by stage and source, not unique test questions. A matching option or passage need not be the same question, and a clean screen cannot rule out semantic or pretraining overlap.
|
| 79 |
+
|
| 80 |
+
We did not produce the planned calibration read or a score without the DI-S selection sample.
|
| 81 |
+
|
| 82 |
+
</details>
|
| 83 |
+
|
| 84 |
### Decision Index 0.1 (archived edition)
|
| 85 |
|
| 86 |
**Full suite: blink-mimo-9b 56.53 vs Jev 1.13.0 59.51.**
|
|
|
|
| 97 |
| Kev 9B | 9B | 50.48 | 32.96 | 30.54 |
|
| 98 |
| Kev 4B | 4B | 47.43 | 28.86 | 25.67 |
|
| 99 |
|
| 100 |
+
We ran the complete archived 0.1 suite: 132,422 requests across 37 benchmarks. The headline index averages 19 panel benchmarks. Comparison rows use the 2026-09-22 leaderboard snapshot. We ran the official kit's scorer locally; these aren't leaderboard submissions. The live [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index) moved to 0.2 on 2026-09-24. Our local 0.2 run is in the section above.
|
|
|
|
|
|
|
| 101 |
|
| 102 |
Third in our local archived 0.1 comparison, behind blink-27b and Jev and ahead of every open entry in the September 22 snapshot (best: Jevfire).
|
| 103 |
|
|
|
|
| 334 |
- **Training overlap.** Public train splits also used by the 0.1 index: ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, ChessBench (searchless_chess training positions; none of the 5,000 test positions), GSM8K (train split; solution-checking items). We also used ANLI and BANKING77 train splits; they're in the 0.1 suite but outside its index, and both are in the 0.2 panel. The audit below reports what was checked and any shared passages.
|
| 335 |
- **Partitions.** Public-source data included training and development partitions.
|
| 336 |
- **Final-mixture audit.** Rechecked every question row (including teacher-written rows) against the complete 0.1 suite (132,422 requests) and JevBench's 231 public items. The checks looked for exact matches of normalised strings of at least 30 characters in any field and shared 13-word passages in each row's question text (instructions, state.question, state.code). Strings or passages seen in 20 or more suite requests were treated as prompt templates and ignored. No public JevBench item matched under these checks; a separate position check found no shared chess positions.
|
| 337 |
+
- **Suite overlap.** 16 BANKING77/VAST training rows share a 13-word passage with 31 suite requests: 23 of VAST's 3,006 and 8 of BANKING77's 3,080. Two VAST training posts are near-duplicates of a test post, but these rows had no exact normalised-text match under the audit. Dropping those requests leaves the index at 56.53 (VAST 0.7805 → 0.7803); BANKING77 is outside the 0.1 index.
|
| 338 |
- **Audit limits.** The 13-word passage check didn't search long-document bodies or option text. Semantic or pretraining overlap can't be ruled out, and private JevBench items weren't available to check.
|
| 339 |
- **Generated reasoning.** Our programs computed the labels for CRUXEval-style code and CLadder-style causal questions; no items from those benchmarks were used. We didn't reuse the suite's GSM8K distractors.
|
| 340 |
- **Teacher documents.** We kept Qwen3.8-27B's documents only if a fresh blind solve by that same teacher agreed with the answer. That's an agreement filter, not independent verification.
|
eval/decision-index-0.2-local.json
ADDED
|
@@ -0,0 +1,1482 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"engine": "blink-mimo-9b",
|
| 3 |
+
"edition": "Decision Index 0.2 (local run, descriptive)",
|
| 4 |
+
"generated_utc": "2026-09-25T23:10:24+00:00",
|
| 5 |
+
"suite": {
|
| 6 |
+
"edition": "release-v2",
|
| 7 |
+
"requests": 121057,
|
| 8 |
+
"scoreable": 120615,
|
| 9 |
+
"excluded": 442,
|
| 10 |
+
"added_requests": 30419,
|
| 11 |
+
"benchmarks": 44,
|
| 12 |
+
"rows_sha256": "b2b56d6fb636837ca469e689087bdbf373dda8de7638aa2da6793e6eda0792d5",
|
| 13 |
+
"added_sha256": "7429f3c9cdddb772c1cfc42bb2a45e8516b0032152b746e6929f1c8b52f4ce89"
|
| 14 |
+
},
|
| 15 |
+
"completed": 151034,
|
| 16 |
+
"complete": true,
|
| 17 |
+
"counts": {
|
| 18 |
+
"ok": 151034
|
| 19 |
+
},
|
| 20 |
+
"latency_ms": {
|
| 21 |
+
"median": 45.8,
|
| 22 |
+
"p95": 1053.9,
|
| 23 |
+
"mean": 214.9
|
| 24 |
+
},
|
| 25 |
+
"decision_index": 43.36,
|
| 26 |
+
"raw_index": 57.24,
|
| 27 |
+
"scores": {
|
| 28 |
+
"balanced_skill": 43.36,
|
| 29 |
+
"balanced_raw": 57.24,
|
| 30 |
+
"breadth_skill": 42.38
|
| 31 |
+
},
|
| 32 |
+
"areas": [
|
| 33 |
+
{
|
| 34 |
+
"id": "knowledge",
|
| 35 |
+
"label": "Knowledge & Reasoning",
|
| 36 |
+
"raw": 0.4824,
|
| 37 |
+
"skill": 0.3319,
|
| 38 |
+
"coverage": 1.0,
|
| 39 |
+
"n": 10,
|
| 40 |
+
"benchmarks": [
|
| 41 |
+
25,
|
| 42 |
+
30,
|
| 43 |
+
31,
|
| 44 |
+
32,
|
| 45 |
+
33,
|
| 46 |
+
43,
|
| 47 |
+
44,
|
| 48 |
+
45,
|
| 49 |
+
57,
|
| 50 |
+
58
|
| 51 |
+
]
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"id": "language",
|
| 55 |
+
"label": "Language Understanding",
|
| 56 |
+
"raw": 0.6719,
|
| 57 |
+
"skill": 0.5405,
|
| 58 |
+
"coverage": 1.0,
|
| 59 |
+
"n": 10,
|
| 60 |
+
"benchmarks": [
|
| 61 |
+
11,
|
| 62 |
+
12,
|
| 63 |
+
28,
|
| 64 |
+
29,
|
| 65 |
+
38,
|
| 66 |
+
39,
|
| 67 |
+
40,
|
| 68 |
+
41,
|
| 69 |
+
42,
|
| 70 |
+
59
|
| 71 |
+
]
|
| 72 |
+
},
|
| 73 |
+
{
|
| 74 |
+
"id": "retrieval",
|
| 75 |
+
"label": "Retrieval & Classification",
|
| 76 |
+
"raw": 0.5536,
|
| 77 |
+
"skill": 0.4048,
|
| 78 |
+
"coverage": 1.0,
|
| 79 |
+
"n": 7,
|
| 80 |
+
"benchmarks": [
|
| 81 |
+
4,
|
| 82 |
+
5,
|
| 83 |
+
10,
|
| 84 |
+
36,
|
| 85 |
+
37,
|
| 86 |
+
56,
|
| 87 |
+
61
|
| 88 |
+
]
|
| 89 |
+
},
|
| 90 |
+
{
|
| 91 |
+
"id": "tools",
|
| 92 |
+
"label": "Tools & Automation",
|
| 93 |
+
"raw": 0.6519,
|
| 94 |
+
"skill": 0.5702,
|
| 95 |
+
"coverage": 1.0,
|
| 96 |
+
"n": 6,
|
| 97 |
+
"benchmarks": [
|
| 98 |
+
1,
|
| 99 |
+
2,
|
| 100 |
+
3,
|
| 101 |
+
6,
|
| 102 |
+
9,
|
| 103 |
+
62
|
| 104 |
+
]
|
| 105 |
+
},
|
| 106 |
+
{
|
| 107 |
+
"id": "arts",
|
| 108 |
+
"label": "Arts & Human Taste",
|
| 109 |
+
"raw": 0.502,
|
| 110 |
+
"skill": 0.3204,
|
| 111 |
+
"coverage": 1.0,
|
| 112 |
+
"n": 7,
|
| 113 |
+
"benchmarks": [
|
| 114 |
+
20,
|
| 115 |
+
21,
|
| 116 |
+
22,
|
| 117 |
+
23,
|
| 118 |
+
48,
|
| 119 |
+
50,
|
| 120 |
+
64
|
| 121 |
+
]
|
| 122 |
+
}
|
| 123 |
+
],
|
| 124 |
+
"index_benchmarks": {
|
| 125 |
+
"1": {
|
| 126 |
+
"raw": 0.8926,
|
| 127 |
+
"skill": 0.855,
|
| 128 |
+
"coverage": 1.0,
|
| 129 |
+
"random": 0.2592,
|
| 130 |
+
"rule": "track",
|
| 131 |
+
"in_index": true,
|
| 132 |
+
"tracks": []
|
| 133 |
+
},
|
| 134 |
+
"2": {
|
| 135 |
+
"raw": 0.4296,
|
| 136 |
+
"skill": 0.3719,
|
| 137 |
+
"coverage": 1.0,
|
| 138 |
+
"random": 0.0918,
|
| 139 |
+
"rule": "track",
|
| 140 |
+
"in_index": true,
|
| 141 |
+
"tracks": []
|
| 142 |
+
},
|
| 143 |
+
"3": {
|
| 144 |
+
"raw": 0.8386,
|
| 145 |
+
"skill": 0.8355,
|
| 146 |
+
"coverage": 1.0,
|
| 147 |
+
"random": 0.0189,
|
| 148 |
+
"rule": "chance",
|
| 149 |
+
"in_index": true
|
| 150 |
+
},
|
| 151 |
+
"4": {
|
| 152 |
+
"raw": 0.8177,
|
| 153 |
+
"skill": 0.8154,
|
| 154 |
+
"coverage": 1.0,
|
| 155 |
+
"random": 0.0127,
|
| 156 |
+
"rule": "chance",
|
| 157 |
+
"in_index": true
|
| 158 |
+
},
|
| 159 |
+
"5": {
|
| 160 |
+
"raw": 0.8381,
|
| 161 |
+
"skill": 0.8371,
|
| 162 |
+
"coverage": 1.0,
|
| 163 |
+
"random": 0.006,
|
| 164 |
+
"rule": "chance",
|
| 165 |
+
"in_index": true
|
| 166 |
+
},
|
| 167 |
+
"6": {
|
| 168 |
+
"raw": 0.7989,
|
| 169 |
+
"skill": 0.5256,
|
| 170 |
+
"coverage": 1.0,
|
| 171 |
+
"random": 0.5245,
|
| 172 |
+
"rule": "track",
|
| 173 |
+
"in_index": true,
|
| 174 |
+
"tracks": [
|
| 175 |
+
{
|
| 176 |
+
"track": "RouterBench-0shot",
|
| 177 |
+
"score": 0.7864,
|
| 178 |
+
"headline": false
|
| 179 |
+
},
|
| 180 |
+
{
|
| 181 |
+
"track": "RouterBench-5shot",
|
| 182 |
+
"score": 0.8114,
|
| 183 |
+
"headline": false
|
| 184 |
+
}
|
| 185 |
+
]
|
| 186 |
+
},
|
| 187 |
+
"9": {
|
| 188 |
+
"raw": 0.3063,
|
| 189 |
+
"skill": 0.3063,
|
| 190 |
+
"coverage": 1.0,
|
| 191 |
+
"random": 0.0,
|
| 192 |
+
"rule": "chance",
|
| 193 |
+
"in_index": true
|
| 194 |
+
},
|
| 195 |
+
"10": {
|
| 196 |
+
"raw": 0.1982,
|
| 197 |
+
"skill": 0.0,
|
| 198 |
+
"coverage": 1.0,
|
| 199 |
+
"random": 0.399,
|
| 200 |
+
"rule": "chance",
|
| 201 |
+
"in_index": true
|
| 202 |
+
},
|
| 203 |
+
"11": {
|
| 204 |
+
"raw": 0.8171,
|
| 205 |
+
"skill": 0.7355,
|
| 206 |
+
"coverage": 1.0,
|
| 207 |
+
"random": 0.3085,
|
| 208 |
+
"rule": "track",
|
| 209 |
+
"in_index": true,
|
| 210 |
+
"tracks": []
|
| 211 |
+
},
|
| 212 |
+
"12": {
|
| 213 |
+
"raw": 0.6899,
|
| 214 |
+
"skill": 0.5355,
|
| 215 |
+
"coverage": 1.0,
|
| 216 |
+
"random": 0.3324,
|
| 217 |
+
"rule": "chance",
|
| 218 |
+
"in_index": true
|
| 219 |
+
},
|
| 220 |
+
"20": {
|
| 221 |
+
"raw": 0.8408,
|
| 222 |
+
"skill": 0.6817,
|
| 223 |
+
"coverage": 1.0,
|
| 224 |
+
"random": 0.5,
|
| 225 |
+
"rule": "track",
|
| 226 |
+
"in_index": true,
|
| 227 |
+
"tracks": []
|
| 228 |
+
},
|
| 229 |
+
"21": {
|
| 230 |
+
"raw": 0.6381,
|
| 231 |
+
"skill": 0.2763,
|
| 232 |
+
"coverage": 1.0,
|
| 233 |
+
"random": 0.5,
|
| 234 |
+
"rule": "track",
|
| 235 |
+
"in_index": true,
|
| 236 |
+
"tracks": []
|
| 237 |
+
},
|
| 238 |
+
"22": {
|
| 239 |
+
"raw": 0.0763,
|
| 240 |
+
"skill": 0.0691,
|
| 241 |
+
"coverage": 1.0,
|
| 242 |
+
"random": 0.0078,
|
| 243 |
+
"rule": "track",
|
| 244 |
+
"in_index": true,
|
| 245 |
+
"tracks": []
|
| 246 |
+
},
|
| 247 |
+
"23": {
|
| 248 |
+
"raw": 0.5965,
|
| 249 |
+
"skill": 0.193,
|
| 250 |
+
"coverage": 1.0,
|
| 251 |
+
"random": 0.5,
|
| 252 |
+
"rule": "track",
|
| 253 |
+
"in_index": true,
|
| 254 |
+
"tracks": []
|
| 255 |
+
},
|
| 256 |
+
"25": {
|
| 257 |
+
"raw": 0.4082,
|
| 258 |
+
"skill": 0.2109,
|
| 259 |
+
"coverage": 1.0,
|
| 260 |
+
"random": 0.25,
|
| 261 |
+
"rule": "track",
|
| 262 |
+
"in_index": true,
|
| 263 |
+
"tracks": []
|
| 264 |
+
},
|
| 265 |
+
"28": {
|
| 266 |
+
"raw": 0.7419,
|
| 267 |
+
"skill": 0.4838,
|
| 268 |
+
"coverage": 1.0,
|
| 269 |
+
"random": 0.5,
|
| 270 |
+
"rule": "chance",
|
| 271 |
+
"in_index": true
|
| 272 |
+
},
|
| 273 |
+
"29": {
|
| 274 |
+
"raw": 0.867,
|
| 275 |
+
"skill": 0.8227,
|
| 276 |
+
"coverage": 1.0,
|
| 277 |
+
"random": 0.25,
|
| 278 |
+
"rule": "chance",
|
| 279 |
+
"in_index": true
|
| 280 |
+
},
|
| 281 |
+
"30": {
|
| 282 |
+
"raw": 0.6585,
|
| 283 |
+
"skill": 0.5882,
|
| 284 |
+
"coverage": 1.0,
|
| 285 |
+
"random": 0.25,
|
| 286 |
+
"rule": "track",
|
| 287 |
+
"in_index": true,
|
| 288 |
+
"tracks": [
|
| 289 |
+
{
|
| 290 |
+
"track": "GSM8K-4choice",
|
| 291 |
+
"score": 0.7096,
|
| 292 |
+
"headline": false
|
| 293 |
+
},
|
| 294 |
+
{
|
| 295 |
+
"track": "GSM8K-10choice",
|
| 296 |
+
"score": 0.6073,
|
| 297 |
+
"headline": false
|
| 298 |
+
}
|
| 299 |
+
]
|
| 300 |
+
},
|
| 301 |
+
"31": {
|
| 302 |
+
"raw": 0.2292,
|
| 303 |
+
"skill": 0.1605,
|
| 304 |
+
"coverage": 1.0,
|
| 305 |
+
"random": 0.0819,
|
| 306 |
+
"rule": "track",
|
| 307 |
+
"in_index": true,
|
| 308 |
+
"tracks": []
|
| 309 |
+
},
|
| 310 |
+
"32": {
|
| 311 |
+
"raw": 0.5705,
|
| 312 |
+
"skill": 0.3172,
|
| 313 |
+
"coverage": 1.0,
|
| 314 |
+
"random": 0.371,
|
| 315 |
+
"rule": "chance",
|
| 316 |
+
"in_index": true
|
| 317 |
+
},
|
| 318 |
+
"33": {
|
| 319 |
+
"raw": 0.3467,
|
| 320 |
+
"skill": 0.338,
|
| 321 |
+
"coverage": 1.0,
|
| 322 |
+
"random": 0.0131,
|
| 323 |
+
"rule": "chance",
|
| 324 |
+
"in_index": true
|
| 325 |
+
},
|
| 326 |
+
"36": {
|
| 327 |
+
"raw": 0.1795,
|
| 328 |
+
"skill": 0.1396,
|
| 329 |
+
"coverage": 1.0,
|
| 330 |
+
"random": 0.0464,
|
| 331 |
+
"rule": "track",
|
| 332 |
+
"in_index": true,
|
| 333 |
+
"tracks": []
|
| 334 |
+
},
|
| 335 |
+
"37": {
|
| 336 |
+
"raw": 0.5199,
|
| 337 |
+
"skill": 0.3978,
|
| 338 |
+
"coverage": 1.0,
|
| 339 |
+
"random": 0.2027,
|
| 340 |
+
"rule": "track",
|
| 341 |
+
"in_index": true,
|
| 342 |
+
"tracks": []
|
| 343 |
+
},
|
| 344 |
+
"38": {
|
| 345 |
+
"raw": 0.055,
|
| 346 |
+
"skill": 0.055,
|
| 347 |
+
"coverage": 1.0,
|
| 348 |
+
"random": 0.0,
|
| 349 |
+
"rule": "chance",
|
| 350 |
+
"in_index": true
|
| 351 |
+
},
|
| 352 |
+
"39": {
|
| 353 |
+
"raw": 0.8246,
|
| 354 |
+
"skill": 0.742,
|
| 355 |
+
"coverage": 1.0,
|
| 356 |
+
"random": 0.3201,
|
| 357 |
+
"rule": "chance",
|
| 358 |
+
"in_index": true
|
| 359 |
+
},
|
| 360 |
+
"40": {
|
| 361 |
+
"raw": 0.5057,
|
| 362 |
+
"skill": 0.3641,
|
| 363 |
+
"coverage": 1.0,
|
| 364 |
+
"random": 0.2227,
|
| 365 |
+
"rule": "track",
|
| 366 |
+
"in_index": true,
|
| 367 |
+
"tracks": [
|
| 368 |
+
{
|
| 369 |
+
"track": "A \u00b7 Arabic",
|
| 370 |
+
"score": 0.3205,
|
| 371 |
+
"headline": false
|
| 372 |
+
},
|
| 373 |
+
{
|
| 374 |
+
"track": "A \u00b7 English",
|
| 375 |
+
"score": 0.5057,
|
| 376 |
+
"headline": true
|
| 377 |
+
},
|
| 378 |
+
{
|
| 379 |
+
"track": "C \u00b7 Arabic pairs",
|
| 380 |
+
"score": 0.79,
|
| 381 |
+
"headline": false
|
| 382 |
+
},
|
| 383 |
+
{
|
| 384 |
+
"track": "C \u00b7 English pairs",
|
| 385 |
+
"score": 0.95,
|
| 386 |
+
"headline": false
|
| 387 |
+
}
|
| 388 |
+
]
|
| 389 |
+
},
|
| 390 |
+
"41": {
|
| 391 |
+
"raw": 0.7805,
|
| 392 |
+
"skill": 0.6707,
|
| 393 |
+
"coverage": 1.0,
|
| 394 |
+
"random": 0.3333,
|
| 395 |
+
"rule": "track",
|
| 396 |
+
"in_index": true,
|
| 397 |
+
"tracks": []
|
| 398 |
+
},
|
| 399 |
+
"42": {
|
| 400 |
+
"raw": 0.8038,
|
| 401 |
+
"skill": 0.6183,
|
| 402 |
+
"coverage": 1.0,
|
| 403 |
+
"random": 0.486,
|
| 404 |
+
"rule": "chance",
|
| 405 |
+
"in_index": true
|
| 406 |
+
},
|
| 407 |
+
"43": {
|
| 408 |
+
"raw": 0.5474,
|
| 409 |
+
"skill": 0.2819,
|
| 410 |
+
"coverage": 1.0,
|
| 411 |
+
"random": 0.3697,
|
| 412 |
+
"rule": "track",
|
| 413 |
+
"in_index": true,
|
| 414 |
+
"tracks": []
|
| 415 |
+
},
|
| 416 |
+
"44": {
|
| 417 |
+
"raw": 0.6614,
|
| 418 |
+
"skill": 0.3228,
|
| 419 |
+
"coverage": 1.0,
|
| 420 |
+
"random": 0.5,
|
| 421 |
+
"rule": "track",
|
| 422 |
+
"in_index": true,
|
| 423 |
+
"tracks": []
|
| 424 |
+
},
|
| 425 |
+
"45": {
|
| 426 |
+
"raw": 0.1098,
|
| 427 |
+
"skill": 0.0,
|
| 428 |
+
"coverage": 1.0,
|
| 429 |
+
"random": 0.1641,
|
| 430 |
+
"rule": "chance",
|
| 431 |
+
"in_index": true
|
| 432 |
+
},
|
| 433 |
+
"48": {
|
| 434 |
+
"raw": 0.282,
|
| 435 |
+
"skill": 0.282,
|
| 436 |
+
"coverage": 1.0,
|
| 437 |
+
"random": 0.25,
|
| 438 |
+
"rule": "vs baseline",
|
| 439 |
+
"in_index": true
|
| 440 |
+
},
|
| 441 |
+
"50": {
|
| 442 |
+
"raw": 0.4553,
|
| 443 |
+
"skill": 0.2094,
|
| 444 |
+
"coverage": 1.0,
|
| 445 |
+
"random": 0.311,
|
| 446 |
+
"rule": "track",
|
| 447 |
+
"in_index": true,
|
| 448 |
+
"tracks": []
|
| 449 |
+
},
|
| 450 |
+
"56": {
|
| 451 |
+
"raw": 0.6655,
|
| 452 |
+
"skill": 0.331,
|
| 453 |
+
"coverage": 1.0,
|
| 454 |
+
"random": 0.5,
|
| 455 |
+
"rule": "chance",
|
| 456 |
+
"in_index": true
|
| 457 |
+
},
|
| 458 |
+
"57": {
|
| 459 |
+
"raw": 0.6137,
|
| 460 |
+
"skill": 0.5655,
|
| 461 |
+
"coverage": 1.0,
|
| 462 |
+
"random": 0.1109,
|
| 463 |
+
"rule": "chance",
|
| 464 |
+
"in_index": true
|
| 465 |
+
},
|
| 466 |
+
"58": {
|
| 467 |
+
"raw": 0.6784,
|
| 468 |
+
"skill": 0.5338,
|
| 469 |
+
"coverage": 1.0,
|
| 470 |
+
"random": 0.3101,
|
| 471 |
+
"rule": "chance",
|
| 472 |
+
"in_index": true
|
| 473 |
+
},
|
| 474 |
+
"59": {
|
| 475 |
+
"raw": 0.6336,
|
| 476 |
+
"skill": 0.3776,
|
| 477 |
+
"coverage": 1.0,
|
| 478 |
+
"random": 0.4113,
|
| 479 |
+
"rule": "chance",
|
| 480 |
+
"in_index": true
|
| 481 |
+
},
|
| 482 |
+
"61": {
|
| 483 |
+
"raw": 0.6565,
|
| 484 |
+
"skill": 0.313,
|
| 485 |
+
"coverage": 1.0,
|
| 486 |
+
"random": 0.5,
|
| 487 |
+
"rule": "chance",
|
| 488 |
+
"in_index": true
|
| 489 |
+
},
|
| 490 |
+
"62": {
|
| 491 |
+
"raw": 0.6451,
|
| 492 |
+
"skill": 0.5268,
|
| 493 |
+
"coverage": 1.0,
|
| 494 |
+
"random": 0.25,
|
| 495 |
+
"rule": "chance",
|
| 496 |
+
"in_index": true
|
| 497 |
+
},
|
| 498 |
+
"64": {
|
| 499 |
+
"raw": 0.625,
|
| 500 |
+
"skill": 0.5312,
|
| 501 |
+
"coverage": 1.0,
|
| 502 |
+
"random": 0.2,
|
| 503 |
+
"rule": "chance",
|
| 504 |
+
"in_index": true
|
| 505 |
+
},
|
| 506 |
+
"24": {
|
| 507 |
+
"raw": 0.788,
|
| 508 |
+
"skill": 0.7173,
|
| 509 |
+
"coverage": 1.0,
|
| 510 |
+
"random": 0.25,
|
| 511 |
+
"rule": "shown, not counted",
|
| 512 |
+
"in_index": false
|
| 513 |
+
},
|
| 514 |
+
"26": {
|
| 515 |
+
"raw": 0.9798,
|
| 516 |
+
"skill": 0.9731,
|
| 517 |
+
"coverage": 1.0,
|
| 518 |
+
"random": 0.2502,
|
| 519 |
+
"rule": "shown, not counted",
|
| 520 |
+
"in_index": false
|
| 521 |
+
},
|
| 522 |
+
"27": {
|
| 523 |
+
"raw": 0.9514,
|
| 524 |
+
"skill": 0.9352,
|
| 525 |
+
"coverage": 1.0,
|
| 526 |
+
"random": 0.2502,
|
| 527 |
+
"rule": "shown, not counted",
|
| 528 |
+
"in_index": false
|
| 529 |
+
},
|
| 530 |
+
"34": {
|
| 531 |
+
"raw": 0.1,
|
| 532 |
+
"skill": 0.0,
|
| 533 |
+
"coverage": 1.0,
|
| 534 |
+
"random": 0.1667,
|
| 535 |
+
"rule": "shown, not counted",
|
| 536 |
+
"in_index": false
|
| 537 |
+
}
|
| 538 |
+
},
|
| 539 |
+
"benchmarks": {
|
| 540 |
+
"1": {
|
| 541 |
+
"catalog_id": 1,
|
| 542 |
+
"dataset": "BFCL",
|
| 543 |
+
"requests": 1694,
|
| 544 |
+
"answered": 1694,
|
| 545 |
+
"unsupported": 0,
|
| 546 |
+
"errors": 0,
|
| 547 |
+
"abstained": 0,
|
| 548 |
+
"pending": 0,
|
| 549 |
+
"scored_requests": 1694,
|
| 550 |
+
"metric": "case exact accuracy",
|
| 551 |
+
"score": 0.8926,
|
| 552 |
+
"reference_same_cases": null,
|
| 553 |
+
"median_ms": 139.8,
|
| 554 |
+
"index_raw": 0.8926,
|
| 555 |
+
"index_skill": 0.855,
|
| 556 |
+
"coverage": 1.0,
|
| 557 |
+
"chance": 0.2592,
|
| 558 |
+
"in_index": true
|
| 559 |
+
},
|
| 560 |
+
"2": {
|
| 561 |
+
"catalog_id": 2,
|
| 562 |
+
"dataset": "ToolRet",
|
| 563 |
+
"requests": 1000,
|
| 564 |
+
"answered": 1000,
|
| 565 |
+
"unsupported": 0,
|
| 566 |
+
"errors": 0,
|
| 567 |
+
"abstained": 0,
|
| 568 |
+
"pending": 0,
|
| 569 |
+
"scored_requests": 1000,
|
| 570 |
+
"metric": "nDCG@10",
|
| 571 |
+
"score": 0.4296,
|
| 572 |
+
"reference_same_cases": null,
|
| 573 |
+
"median_ms": 1187.9,
|
| 574 |
+
"index_raw": 0.4296,
|
| 575 |
+
"index_skill": 0.3719,
|
| 576 |
+
"coverage": 1.0,
|
| 577 |
+
"chance": 0.0918,
|
| 578 |
+
"in_index": true
|
| 579 |
+
},
|
| 580 |
+
"3": {
|
| 581 |
+
"catalog_id": 3,
|
| 582 |
+
"dataset": "API-Bank",
|
| 583 |
+
"requests": 508,
|
| 584 |
+
"answered": 508,
|
| 585 |
+
"unsupported": 0,
|
| 586 |
+
"errors": 0,
|
| 587 |
+
"abstained": 0,
|
| 588 |
+
"pending": 0,
|
| 589 |
+
"scored_requests": 508,
|
| 590 |
+
"metric": "accuracy",
|
| 591 |
+
"score": 0.8386,
|
| 592 |
+
"reference_same_cases": null,
|
| 593 |
+
"median_ms": 640.5,
|
| 594 |
+
"index_raw": 0.8386,
|
| 595 |
+
"index_skill": 0.8355,
|
| 596 |
+
"coverage": 1.0,
|
| 597 |
+
"chance": 0.0189,
|
| 598 |
+
"in_index": true
|
| 599 |
+
},
|
| 600 |
+
"4": {
|
| 601 |
+
"catalog_id": 4,
|
| 602 |
+
"dataset": "BANKING77",
|
| 603 |
+
"requests": 3080,
|
| 604 |
+
"answered": 3080,
|
| 605 |
+
"unsupported": 0,
|
| 606 |
+
"errors": 0,
|
| 607 |
+
"abstained": 0,
|
| 608 |
+
"pending": 0,
|
| 609 |
+
"scored_requests": 3080,
|
| 610 |
+
"metric": "macro-F1",
|
| 611 |
+
"score": 0.8177,
|
| 612 |
+
"reference_same_cases": null,
|
| 613 |
+
"median_ms": 103.1,
|
| 614 |
+
"index_raw": 0.8177,
|
| 615 |
+
"index_skill": 0.8154,
|
| 616 |
+
"coverage": 1.0,
|
| 617 |
+
"chance": 0.0127,
|
| 618 |
+
"in_index": true
|
| 619 |
+
},
|
| 620 |
+
"5": {
|
| 621 |
+
"catalog_id": 5,
|
| 622 |
+
"dataset": "CLINC150+OOS",
|
| 623 |
+
"requests": 5500,
|
| 624 |
+
"answered": 5500,
|
| 625 |
+
"unsupported": 0,
|
| 626 |
+
"errors": 0,
|
| 627 |
+
"abstained": 0,
|
| 628 |
+
"pending": 0,
|
| 629 |
+
"scored_requests": 5500,
|
| 630 |
+
"metric": "macro-F1",
|
| 631 |
+
"score": 0.8381,
|
| 632 |
+
"reference_same_cases": null,
|
| 633 |
+
"median_ms": 172.3,
|
| 634 |
+
"index_raw": 0.8381,
|
| 635 |
+
"index_skill": 0.8371,
|
| 636 |
+
"coverage": 1.0,
|
| 637 |
+
"chance": 0.006,
|
| 638 |
+
"in_index": true
|
| 639 |
+
},
|
| 640 |
+
"6": {
|
| 641 |
+
"catalog_id": 6,
|
| 642 |
+
"dataset": "RouterBench",
|
| 643 |
+
"requests": 10000,
|
| 644 |
+
"answered": 10000,
|
| 645 |
+
"unsupported": 0,
|
| 646 |
+
"errors": 0,
|
| 647 |
+
"abstained": 0,
|
| 648 |
+
"pending": 0,
|
| 649 |
+
"scored_requests": 10000,
|
| 650 |
+
"metric": "selected quality (quality objective)",
|
| 651 |
+
"score": 0.7989,
|
| 652 |
+
"reference_same_cases": null,
|
| 653 |
+
"median_ms": 392.5,
|
| 654 |
+
"tracks": [
|
| 655 |
+
{
|
| 656 |
+
"track": "RouterBench-0shot",
|
| 657 |
+
"score": 0.7864,
|
| 658 |
+
"headline": false
|
| 659 |
+
},
|
| 660 |
+
{
|
| 661 |
+
"track": "RouterBench-5shot",
|
| 662 |
+
"score": 0.8114,
|
| 663 |
+
"headline": false
|
| 664 |
+
}
|
| 665 |
+
],
|
| 666 |
+
"index_raw": 0.7989,
|
| 667 |
+
"index_skill": 0.5256,
|
| 668 |
+
"coverage": 1.0,
|
| 669 |
+
"chance": 0.5245,
|
| 670 |
+
"in_index": true
|
| 671 |
+
},
|
| 672 |
+
"9": {
|
| 673 |
+
"catalog_id": 9,
|
| 674 |
+
"dataset": "Home appliance simulator",
|
| 675 |
+
"requests": 160,
|
| 676 |
+
"answered": 160,
|
| 677 |
+
"unsupported": 0,
|
| 678 |
+
"errors": 0,
|
| 679 |
+
"abstained": 0,
|
| 680 |
+
"pending": 0,
|
| 681 |
+
"scored_requests": 160,
|
| 682 |
+
"metric": "case exact accuracy",
|
| 683 |
+
"score": 0.3063,
|
| 684 |
+
"reference_same_cases": null,
|
| 685 |
+
"median_ms": 1713.5,
|
| 686 |
+
"index_raw": 0.3063,
|
| 687 |
+
"index_skill": 0.3063,
|
| 688 |
+
"coverage": 1.0,
|
| 689 |
+
"chance": 0.0,
|
| 690 |
+
"in_index": true
|
| 691 |
+
},
|
| 692 |
+
"10": {
|
| 693 |
+
"catalog_id": 10,
|
| 694 |
+
"dataset": "SGD/SGD-X",
|
| 695 |
+
"requests": 2500,
|
| 696 |
+
"answered": 2500,
|
| 697 |
+
"unsupported": 0,
|
| 698 |
+
"errors": 0,
|
| 699 |
+
"abstained": 0,
|
| 700 |
+
"pending": 0,
|
| 701 |
+
"scored_requests": 2500,
|
| 702 |
+
"metric": "macro-F1",
|
| 703 |
+
"score": 0.1982,
|
| 704 |
+
"reference_same_cases": null,
|
| 705 |
+
"median_ms": 106.5,
|
| 706 |
+
"index_raw": 0.1982,
|
| 707 |
+
"index_skill": 0.0,
|
| 708 |
+
"coverage": 1.0,
|
| 709 |
+
"chance": 0.399,
|
| 710 |
+
"in_index": true
|
| 711 |
+
},
|
| 712 |
+
"11": {
|
| 713 |
+
"catalog_id": 11,
|
| 714 |
+
"dataset": "ContractNLI",
|
| 715 |
+
"requests": 123,
|
| 716 |
+
"answered": 123,
|
| 717 |
+
"unsupported": 0,
|
| 718 |
+
"errors": 0,
|
| 719 |
+
"abstained": 0,
|
| 720 |
+
"pending": 0,
|
| 721 |
+
"scored_requests": 123,
|
| 722 |
+
"metric": "macro-F1",
|
| 723 |
+
"score": 0.8171,
|
| 724 |
+
"reference_same_cases": null,
|
| 725 |
+
"median_ms": 2966.1,
|
| 726 |
+
"index_raw": 0.8171,
|
| 727 |
+
"index_skill": 0.7355,
|
| 728 |
+
"coverage": 1.0,
|
| 729 |
+
"chance": 0.3085,
|
| 730 |
+
"in_index": true
|
| 731 |
+
},
|
| 732 |
+
"12": {
|
| 733 |
+
"catalog_id": 12,
|
| 734 |
+
"dataset": "ANLI",
|
| 735 |
+
"requests": 3200,
|
| 736 |
+
"answered": 3200,
|
| 737 |
+
"unsupported": 0,
|
| 738 |
+
"errors": 0,
|
| 739 |
+
"abstained": 0,
|
| 740 |
+
"pending": 0,
|
| 741 |
+
"scored_requests": 3200,
|
| 742 |
+
"metric": "macro-F1",
|
| 743 |
+
"score": 0.6899,
|
| 744 |
+
"reference_same_cases": null,
|
| 745 |
+
"median_ms": 19.4,
|
| 746 |
+
"index_raw": 0.6899,
|
| 747 |
+
"index_skill": 0.5355,
|
| 748 |
+
"coverage": 1.0,
|
| 749 |
+
"chance": 0.3324,
|
| 750 |
+
"in_index": true
|
| 751 |
+
},
|
| 752 |
+
"20": {
|
| 753 |
+
"catalog_id": 20,
|
| 754 |
+
"dataset": "BPoMP",
|
| 755 |
+
"requests": 5000,
|
| 756 |
+
"answered": 5000,
|
| 757 |
+
"unsupported": 0,
|
| 758 |
+
"errors": 0,
|
| 759 |
+
"abstained": 0,
|
| 760 |
+
"pending": 0,
|
| 761 |
+
"scored_requests": 5000,
|
| 762 |
+
"metric": "accuracy",
|
| 763 |
+
"score": 0.8398,
|
| 764 |
+
"reference_same_cases": null,
|
| 765 |
+
"median_ms": 18.7,
|
| 766 |
+
"index_raw": 0.8408,
|
| 767 |
+
"index_skill": 0.6817,
|
| 768 |
+
"coverage": 1.0,
|
| 769 |
+
"chance": 0.5,
|
| 770 |
+
"in_index": true
|
| 771 |
+
},
|
| 772 |
+
"21": {
|
| 773 |
+
"catalog_id": 21,
|
| 774 |
+
"dataset": "Humicroedit",
|
| 775 |
+
"requests": 2628,
|
| 776 |
+
"answered": 2628,
|
| 777 |
+
"unsupported": 0,
|
| 778 |
+
"errors": 0,
|
| 779 |
+
"abstained": 0,
|
| 780 |
+
"pending": 0,
|
| 781 |
+
"scored_requests": 2628,
|
| 782 |
+
"metric": "accuracy",
|
| 783 |
+
"score": 0.6381,
|
| 784 |
+
"reference_same_cases": null,
|
| 785 |
+
"median_ms": 11.8,
|
| 786 |
+
"index_raw": 0.6381,
|
| 787 |
+
"index_skill": 0.2763,
|
| 788 |
+
"coverage": 1.0,
|
| 789 |
+
"chance": 0.5,
|
| 790 |
+
"in_index": true
|
| 791 |
+
},
|
| 792 |
+
"22": {
|
| 793 |
+
"catalog_id": 22,
|
| 794 |
+
"dataset": "POP909-CL",
|
| 795 |
+
"requests": 2000,
|
| 796 |
+
"answered": 2000,
|
| 797 |
+
"unsupported": 0,
|
| 798 |
+
"errors": 0,
|
| 799 |
+
"abstained": 0,
|
| 800 |
+
"pending": 0,
|
| 801 |
+
"scored_requests": 2000,
|
| 802 |
+
"metric": "accuracy",
|
| 803 |
+
"score": 0.0825,
|
| 804 |
+
"reference_same_cases": null,
|
| 805 |
+
"median_ms": 843.5,
|
| 806 |
+
"index_raw": 0.0763,
|
| 807 |
+
"index_skill": 0.0691,
|
| 808 |
+
"coverage": 1.0,
|
| 809 |
+
"chance": 0.0078,
|
| 810 |
+
"in_index": true
|
| 811 |
+
},
|
| 812 |
+
"23": {
|
| 813 |
+
"catalog_id": 23,
|
| 814 |
+
"dataset": "cfcolor",
|
| 815 |
+
"requests": 5000,
|
| 816 |
+
"answered": 5000,
|
| 817 |
+
"unsupported": 0,
|
| 818 |
+
"errors": 0,
|
| 819 |
+
"abstained": 0,
|
| 820 |
+
"pending": 0,
|
| 821 |
+
"scored_requests": 5000,
|
| 822 |
+
"metric": "accuracy",
|
| 823 |
+
"score": 0.5886,
|
| 824 |
+
"reference_same_cases": null,
|
| 825 |
+
"median_ms": 50.0,
|
| 826 |
+
"index_raw": 0.5965,
|
| 827 |
+
"index_skill": 0.193,
|
| 828 |
+
"coverage": 1.0,
|
| 829 |
+
"chance": 0.5,
|
| 830 |
+
"in_index": true
|
| 831 |
+
},
|
| 832 |
+
"24": {
|
| 833 |
+
"catalog_id": 24,
|
| 834 |
+
"dataset": "MMLU",
|
| 835 |
+
"requests": 14033,
|
| 836 |
+
"answered": 14033,
|
| 837 |
+
"unsupported": 0,
|
| 838 |
+
"errors": 0,
|
| 839 |
+
"abstained": 0,
|
| 840 |
+
"pending": 0,
|
| 841 |
+
"scored_requests": 14033,
|
| 842 |
+
"metric": "accuracy",
|
| 843 |
+
"score": 0.788,
|
| 844 |
+
"reference_same_cases": null,
|
| 845 |
+
"median_ms": 18.2,
|
| 846 |
+
"index_raw": 0.788,
|
| 847 |
+
"index_skill": 0.7173,
|
| 848 |
+
"coverage": 1.0,
|
| 849 |
+
"chance": 0.25,
|
| 850 |
+
"in_index": false
|
| 851 |
+
},
|
| 852 |
+
"25": {
|
| 853 |
+
"catalog_id": 25,
|
| 854 |
+
"dataset": "GPQA Diamond",
|
| 855 |
+
"requests": 196,
|
| 856 |
+
"answered": 196,
|
| 857 |
+
"unsupported": 0,
|
| 858 |
+
"errors": 0,
|
| 859 |
+
"abstained": 0,
|
| 860 |
+
"pending": 0,
|
| 861 |
+
"scored_requests": 196,
|
| 862 |
+
"metric": "accuracy",
|
| 863 |
+
"score": 0.4082,
|
| 864 |
+
"reference_same_cases": null,
|
| 865 |
+
"median_ms": 34.6,
|
| 866 |
+
"index_raw": 0.4082,
|
| 867 |
+
"index_skill": 0.2109,
|
| 868 |
+
"coverage": 1.0,
|
| 869 |
+
"chance": 0.25,
|
| 870 |
+
"in_index": true
|
| 871 |
+
},
|
| 872 |
+
"26": {
|
| 873 |
+
"catalog_id": 26,
|
| 874 |
+
"dataset": "ARC-Easy",
|
| 875 |
+
"requests": 2376,
|
| 876 |
+
"answered": 2376,
|
| 877 |
+
"unsupported": 0,
|
| 878 |
+
"errors": 0,
|
| 879 |
+
"abstained": 0,
|
| 880 |
+
"pending": 0,
|
| 881 |
+
"scored_requests": 2376,
|
| 882 |
+
"metric": "accuracy",
|
| 883 |
+
"score": 0.9798,
|
| 884 |
+
"reference_same_cases": null,
|
| 885 |
+
"median_ms": 16.1,
|
| 886 |
+
"index_raw": 0.9798,
|
| 887 |
+
"index_skill": 0.9731,
|
| 888 |
+
"coverage": 1.0,
|
| 889 |
+
"chance": 0.2502,
|
| 890 |
+
"in_index": false
|
| 891 |
+
},
|
| 892 |
+
"27": {
|
| 893 |
+
"catalog_id": 27,
|
| 894 |
+
"dataset": "ARC-Challenge",
|
| 895 |
+
"requests": 1172,
|
| 896 |
+
"answered": 1172,
|
| 897 |
+
"unsupported": 0,
|
| 898 |
+
"errors": 0,
|
| 899 |
+
"abstained": 0,
|
| 900 |
+
"pending": 0,
|
| 901 |
+
"scored_requests": 1172,
|
| 902 |
+
"metric": "accuracy",
|
| 903 |
+
"score": 0.9514,
|
| 904 |
+
"reference_same_cases": null,
|
| 905 |
+
"median_ms": 16.9,
|
| 906 |
+
"index_raw": 0.9514,
|
| 907 |
+
"index_skill": 0.9352,
|
| 908 |
+
"coverage": 1.0,
|
| 909 |
+
"chance": 0.2502,
|
| 910 |
+
"in_index": false
|
| 911 |
+
},
|
| 912 |
+
"28": {
|
| 913 |
+
"catalog_id": 28,
|
| 914 |
+
"dataset": "WinoGrande",
|
| 915 |
+
"requests": 1267,
|
| 916 |
+
"answered": 1267,
|
| 917 |
+
"unsupported": 0,
|
| 918 |
+
"errors": 0,
|
| 919 |
+
"abstained": 0,
|
| 920 |
+
"pending": 0,
|
| 921 |
+
"scored_requests": 1267,
|
| 922 |
+
"metric": "accuracy",
|
| 923 |
+
"score": 0.7419,
|
| 924 |
+
"reference_same_cases": null,
|
| 925 |
+
"median_ms": 12.0,
|
| 926 |
+
"index_raw": 0.7419,
|
| 927 |
+
"index_skill": 0.4838,
|
| 928 |
+
"coverage": 1.0,
|
| 929 |
+
"chance": 0.5,
|
| 930 |
+
"in_index": true
|
| 931 |
+
},
|
| 932 |
+
"29": {
|
| 933 |
+
"catalog_id": 29,
|
| 934 |
+
"dataset": "HellaSwag",
|
| 935 |
+
"requests": 10042,
|
| 936 |
+
"answered": 10042,
|
| 937 |
+
"unsupported": 0,
|
| 938 |
+
"errors": 0,
|
| 939 |
+
"abstained": 0,
|
| 940 |
+
"pending": 0,
|
| 941 |
+
"scored_requests": 10042,
|
| 942 |
+
"metric": "accuracy",
|
| 943 |
+
"score": 0.867,
|
| 944 |
+
"reference_same_cases": null,
|
| 945 |
+
"median_ms": 27.6,
|
| 946 |
+
"index_raw": 0.867,
|
| 947 |
+
"index_skill": 0.8227,
|
| 948 |
+
"coverage": 1.0,
|
| 949 |
+
"chance": 0.25,
|
| 950 |
+
"in_index": true
|
| 951 |
+
},
|
| 952 |
+
"30": {
|
| 953 |
+
"catalog_id": 30,
|
| 954 |
+
"dataset": "GSM8K",
|
| 955 |
+
"requests": 2638,
|
| 956 |
+
"answered": 2638,
|
| 957 |
+
"unsupported": 0,
|
| 958 |
+
"errors": 0,
|
| 959 |
+
"abstained": 0,
|
| 960 |
+
"pending": 0,
|
| 961 |
+
"scored_requests": 2638,
|
| 962 |
+
"metric": "accuracy",
|
| 963 |
+
"score": 0.6585,
|
| 964 |
+
"reference_same_cases": null,
|
| 965 |
+
"median_ms": 24.3,
|
| 966 |
+
"tracks": [
|
| 967 |
+
{
|
| 968 |
+
"track": "GSM8K-4choice",
|
| 969 |
+
"score": 0.7096,
|
| 970 |
+
"headline": false
|
| 971 |
+
},
|
| 972 |
+
{
|
| 973 |
+
"track": "GSM8K-10choice",
|
| 974 |
+
"score": 0.6073,
|
| 975 |
+
"headline": false
|
| 976 |
+
}
|
| 977 |
+
],
|
| 978 |
+
"index_raw": 0.6585,
|
| 979 |
+
"index_skill": 0.5882,
|
| 980 |
+
"coverage": 1.0,
|
| 981 |
+
"chance": 0.25,
|
| 982 |
+
"in_index": true
|
| 983 |
+
},
|
| 984 |
+
"31": {
|
| 985 |
+
"catalog_id": 31,
|
| 986 |
+
"dataset": "ChessBench",
|
| 987 |
+
"requests": 5000,
|
| 988 |
+
"answered": 5000,
|
| 989 |
+
"unsupported": 0,
|
| 990 |
+
"errors": 0,
|
| 991 |
+
"abstained": 0,
|
| 992 |
+
"pending": 0,
|
| 993 |
+
"scored_requests": 5000,
|
| 994 |
+
"metric": "accuracy",
|
| 995 |
+
"score": 0.2292,
|
| 996 |
+
"reference_same_cases": null,
|
| 997 |
+
"median_ms": 132.4,
|
| 998 |
+
"index_raw": 0.2292,
|
| 999 |
+
"index_skill": 0.1605,
|
| 1000 |
+
"coverage": 1.0,
|
| 1001 |
+
"chance": 0.0819,
|
| 1002 |
+
"in_index": true
|
| 1003 |
+
},
|
| 1004 |
+
"32": {
|
| 1005 |
+
"catalog_id": 32,
|
| 1006 |
+
"dataset": "MuSR",
|
| 1007 |
+
"requests": 752,
|
| 1008 |
+
"answered": 752,
|
| 1009 |
+
"unsupported": 0,
|
| 1010 |
+
"errors": 0,
|
| 1011 |
+
"abstained": 0,
|
| 1012 |
+
"pending": 0,
|
| 1013 |
+
"scored_requests": 752,
|
| 1014 |
+
"metric": "accuracy",
|
| 1015 |
+
"score": 0.5705,
|
| 1016 |
+
"reference_same_cases": null,
|
| 1017 |
+
"median_ms": 93.7,
|
| 1018 |
+
"index_raw": 0.5705,
|
| 1019 |
+
"index_skill": 0.3172,
|
| 1020 |
+
"coverage": 1.0,
|
| 1021 |
+
"chance": 0.371,
|
| 1022 |
+
"in_index": true
|
| 1023 |
+
},
|
| 1024 |
+
"33": {
|
| 1025 |
+
"catalog_id": 33,
|
| 1026 |
+
"dataset": "SATA-Bench",
|
| 1027 |
+
"requests": 1650,
|
| 1028 |
+
"answered": 1650,
|
| 1029 |
+
"unsupported": 0,
|
| 1030 |
+
"errors": 0,
|
| 1031 |
+
"abstained": 0,
|
| 1032 |
+
"pending": 0,
|
| 1033 |
+
"scored_requests": 1650,
|
| 1034 |
+
"metric": "case exact accuracy",
|
| 1035 |
+
"score": 0.3467,
|
| 1036 |
+
"reference_same_cases": null,
|
| 1037 |
+
"median_ms": 321.2,
|
| 1038 |
+
"index_raw": 0.3467,
|
| 1039 |
+
"index_skill": 0.338,
|
| 1040 |
+
"coverage": 1.0,
|
| 1041 |
+
"chance": 0.0131,
|
| 1042 |
+
"in_index": true
|
| 1043 |
+
},
|
| 1044 |
+
"34": {
|
| 1045 |
+
"catalog_id": 34,
|
| 1046 |
+
"dataset": "SimpleBench",
|
| 1047 |
+
"requests": 10,
|
| 1048 |
+
"answered": 10,
|
| 1049 |
+
"unsupported": 0,
|
| 1050 |
+
"errors": 0,
|
| 1051 |
+
"abstained": 0,
|
| 1052 |
+
"pending": 0,
|
| 1053 |
+
"scored_requests": 10,
|
| 1054 |
+
"metric": "accuracy",
|
| 1055 |
+
"score": 0.1,
|
| 1056 |
+
"reference_same_cases": null,
|
| 1057 |
+
"median_ms": 21.2,
|
| 1058 |
+
"index_raw": 0.1,
|
| 1059 |
+
"index_skill": 0.0,
|
| 1060 |
+
"coverage": 1.0,
|
| 1061 |
+
"chance": 0.1667,
|
| 1062 |
+
"in_index": false
|
| 1063 |
+
},
|
| 1064 |
+
"36": {
|
| 1065 |
+
"catalog_id": 36,
|
| 1066 |
+
"dataset": "BRIGHT",
|
| 1067 |
+
"requests": 550,
|
| 1068 |
+
"answered": 550,
|
| 1069 |
+
"unsupported": 0,
|
| 1070 |
+
"errors": 0,
|
| 1071 |
+
"abstained": 0,
|
| 1072 |
+
"pending": 0,
|
| 1073 |
+
"scored_requests": 550,
|
| 1074 |
+
"metric": "nDCG@10",
|
| 1075 |
+
"score": 0.1795,
|
| 1076 |
+
"reference_same_cases": null,
|
| 1077 |
+
"median_ms": 1508.7,
|
| 1078 |
+
"index_raw": 0.1795,
|
| 1079 |
+
"index_skill": 0.1396,
|
| 1080 |
+
"coverage": 1.0,
|
| 1081 |
+
"chance": 0.0464,
|
| 1082 |
+
"in_index": true
|
| 1083 |
+
},
|
| 1084 |
+
"37": {
|
| 1085 |
+
"catalog_id": 37,
|
| 1086 |
+
"dataset": "Amazon ESCI",
|
| 1087 |
+
"requests": 5000,
|
| 1088 |
+
"answered": 5000,
|
| 1089 |
+
"unsupported": 0,
|
| 1090 |
+
"errors": 0,
|
| 1091 |
+
"abstained": 0,
|
| 1092 |
+
"pending": 0,
|
| 1093 |
+
"scored_requests": 5000,
|
| 1094 |
+
"metric": "macro-F1",
|
| 1095 |
+
"score": 0.5199,
|
| 1096 |
+
"reference_same_cases": null,
|
| 1097 |
+
"median_ms": 39.0,
|
| 1098 |
+
"index_raw": 0.5199,
|
| 1099 |
+
"index_skill": 0.3978,
|
| 1100 |
+
"coverage": 1.0,
|
| 1101 |
+
"chance": 0.2027,
|
| 1102 |
+
"in_index": true
|
| 1103 |
+
},
|
| 1104 |
+
"38": {
|
| 1105 |
+
"catalog_id": 38,
|
| 1106 |
+
"dataset": "ACOS",
|
| 1107 |
+
"requests": 1565,
|
| 1108 |
+
"answered": 1565,
|
| 1109 |
+
"unsupported": 0,
|
| 1110 |
+
"errors": 0,
|
| 1111 |
+
"abstained": 0,
|
| 1112 |
+
"pending": 0,
|
| 1113 |
+
"scored_requests": 1565,
|
| 1114 |
+
"metric": "case exact accuracy",
|
| 1115 |
+
"score": 0.055,
|
| 1116 |
+
"reference_same_cases": null,
|
| 1117 |
+
"median_ms": 982.6,
|
| 1118 |
+
"index_raw": 0.055,
|
| 1119 |
+
"index_skill": 0.055,
|
| 1120 |
+
"coverage": 1.0,
|
| 1121 |
+
"chance": 0.0,
|
| 1122 |
+
"in_index": true
|
| 1123 |
+
},
|
| 1124 |
+
"39": {
|
| 1125 |
+
"catalog_id": 39,
|
| 1126 |
+
"dataset": "FinEntity",
|
| 1127 |
+
"requests": 979,
|
| 1128 |
+
"answered": 979,
|
| 1129 |
+
"unsupported": 0,
|
| 1130 |
+
"errors": 0,
|
| 1131 |
+
"abstained": 0,
|
| 1132 |
+
"pending": 0,
|
| 1133 |
+
"scored_requests": 979,
|
| 1134 |
+
"metric": "macro-F1",
|
| 1135 |
+
"score": 0.8246,
|
| 1136 |
+
"reference_same_cases": null,
|
| 1137 |
+
"median_ms": 36.0,
|
| 1138 |
+
"index_raw": 0.8246,
|
| 1139 |
+
"index_skill": 0.742,
|
| 1140 |
+
"coverage": 1.0,
|
| 1141 |
+
"chance": 0.3201,
|
| 1142 |
+
"in_index": true
|
| 1143 |
+
},
|
| 1144 |
+
"40": {
|
| 1145 |
+
"catalog_id": 40,
|
| 1146 |
+
"dataset": "iSarcasmEval",
|
| 1147 |
+
"requests": 4600,
|
| 1148 |
+
"answered": 4600,
|
| 1149 |
+
"unsupported": 0,
|
| 1150 |
+
"errors": 0,
|
| 1151 |
+
"abstained": 0,
|
| 1152 |
+
"pending": 0,
|
| 1153 |
+
"scored_requests": 4600,
|
| 1154 |
+
"metric": "Sarcasm F1 \u00b7 track A, English",
|
| 1155 |
+
"score": 0.5057,
|
| 1156 |
+
"reference_same_cases": null,
|
| 1157 |
+
"median_ms": 12.9,
|
| 1158 |
+
"tracks": [
|
| 1159 |
+
{
|
| 1160 |
+
"track": "A \u00b7 Arabic",
|
| 1161 |
+
"score": 0.3205,
|
| 1162 |
+
"headline": false
|
| 1163 |
+
},
|
| 1164 |
+
{
|
| 1165 |
+
"track": "A \u00b7 English",
|
| 1166 |
+
"score": 0.5057,
|
| 1167 |
+
"headline": true
|
| 1168 |
+
},
|
| 1169 |
+
{
|
| 1170 |
+
"track": "C \u00b7 Arabic pairs",
|
| 1171 |
+
"score": 0.79,
|
| 1172 |
+
"headline": false
|
| 1173 |
+
},
|
| 1174 |
+
{
|
| 1175 |
+
"track": "C \u00b7 English pairs",
|
| 1176 |
+
"score": 0.95,
|
| 1177 |
+
"headline": false
|
| 1178 |
+
}
|
| 1179 |
+
],
|
| 1180 |
+
"index_raw": 0.5057,
|
| 1181 |
+
"index_skill": 0.3641,
|
| 1182 |
+
"coverage": 1.0,
|
| 1183 |
+
"chance": 0.2227,
|
| 1184 |
+
"in_index": true
|
| 1185 |
+
},
|
| 1186 |
+
"41": {
|
| 1187 |
+
"catalog_id": 41,
|
| 1188 |
+
"dataset": "VAST",
|
| 1189 |
+
"requests": 3006,
|
| 1190 |
+
"answered": 3006,
|
| 1191 |
+
"unsupported": 0,
|
| 1192 |
+
"errors": 0,
|
| 1193 |
+
"abstained": 0,
|
| 1194 |
+
"pending": 0,
|
| 1195 |
+
"scored_requests": 3006,
|
| 1196 |
+
"metric": "macro-F1",
|
| 1197 |
+
"score": 0.7805,
|
| 1198 |
+
"reference_same_cases": null,
|
| 1199 |
+
"median_ms": 23.6,
|
| 1200 |
+
"index_raw": 0.7805,
|
| 1201 |
+
"index_skill": 0.6707,
|
| 1202 |
+
"coverage": 1.0,
|
| 1203 |
+
"chance": 0.3333,
|
| 1204 |
+
"in_index": true
|
| 1205 |
+
},
|
| 1206 |
+
"42": {
|
| 1207 |
+
"catalog_id": 42,
|
| 1208 |
+
"dataset": "NLI4CT",
|
| 1209 |
+
"requests": 5500,
|
| 1210 |
+
"answered": 5500,
|
| 1211 |
+
"unsupported": 0,
|
| 1212 |
+
"errors": 0,
|
| 1213 |
+
"abstained": 0,
|
| 1214 |
+
"pending": 0,
|
| 1215 |
+
"scored_requests": 5500,
|
| 1216 |
+
"metric": "macro-F1",
|
| 1217 |
+
"score": 0.8038,
|
| 1218 |
+
"reference_same_cases": null,
|
| 1219 |
+
"median_ms": 54.9,
|
| 1220 |
+
"index_raw": 0.8038,
|
| 1221 |
+
"index_skill": 0.6183,
|
| 1222 |
+
"coverage": 1.0,
|
| 1223 |
+
"chance": 0.486,
|
| 1224 |
+
"in_index": true
|
| 1225 |
+
},
|
| 1226 |
+
"43": {
|
| 1227 |
+
"catalog_id": 43,
|
| 1228 |
+
"dataset": "CRUXEval",
|
| 1229 |
+
"requests": 570,
|
| 1230 |
+
"answered": 570,
|
| 1231 |
+
"unsupported": 0,
|
| 1232 |
+
"errors": 0,
|
| 1233 |
+
"abstained": 0,
|
| 1234 |
+
"pending": 0,
|
| 1235 |
+
"scored_requests": 570,
|
| 1236 |
+
"metric": "accuracy",
|
| 1237 |
+
"score": 0.5474,
|
| 1238 |
+
"reference_same_cases": null,
|
| 1239 |
+
"median_ms": 20.4,
|
| 1240 |
+
"index_raw": 0.5474,
|
| 1241 |
+
"index_skill": 0.2819,
|
| 1242 |
+
"coverage": 1.0,
|
| 1243 |
+
"chance": 0.3697,
|
| 1244 |
+
"in_index": true
|
| 1245 |
+
},
|
| 1246 |
+
"44": {
|
| 1247 |
+
"catalog_id": 44,
|
| 1248 |
+
"dataset": "CLadder",
|
| 1249 |
+
"requests": 5000,
|
| 1250 |
+
"answered": 5000,
|
| 1251 |
+
"unsupported": 0,
|
| 1252 |
+
"errors": 0,
|
| 1253 |
+
"abstained": 0,
|
| 1254 |
+
"pending": 0,
|
| 1255 |
+
"scored_requests": 5000,
|
| 1256 |
+
"metric": "accuracy",
|
| 1257 |
+
"score": 0.6614,
|
| 1258 |
+
"reference_same_cases": null,
|
| 1259 |
+
"median_ms": 19.9,
|
| 1260 |
+
"index_raw": 0.6614,
|
| 1261 |
+
"index_skill": 0.3228,
|
| 1262 |
+
"coverage": 1.0,
|
| 1263 |
+
"chance": 0.5,
|
| 1264 |
+
"in_index": true
|
| 1265 |
+
},
|
| 1266 |
+
"45": {
|
| 1267 |
+
"catalog_id": 45,
|
| 1268 |
+
"dataset": "HLE",
|
| 1269 |
+
"requests": 501,
|
| 1270 |
+
"answered": 501,
|
| 1271 |
+
"unsupported": 0,
|
| 1272 |
+
"errors": 0,
|
| 1273 |
+
"abstained": 0,
|
| 1274 |
+
"pending": 0,
|
| 1275 |
+
"scored_requests": 501,
|
| 1276 |
+
"metric": "accuracy",
|
| 1277 |
+
"score": 0.1098,
|
| 1278 |
+
"reference_same_cases": null,
|
| 1279 |
+
"median_ms": 30.4,
|
| 1280 |
+
"index_raw": 0.1098,
|
| 1281 |
+
"index_skill": 0.0,
|
| 1282 |
+
"coverage": 1.0,
|
| 1283 |
+
"chance": 0.1641,
|
| 1284 |
+
"in_index": true
|
| 1285 |
+
},
|
| 1286 |
+
"48": {
|
| 1287 |
+
"catalog_id": 48,
|
| 1288 |
+
"dataset": "ForecastBench",
|
| 1289 |
+
"requests": 10139,
|
| 1290 |
+
"answered": 10139,
|
| 1291 |
+
"unsupported": 0,
|
| 1292 |
+
"errors": 0,
|
| 1293 |
+
"abstained": 0,
|
| 1294 |
+
"pending": 0,
|
| 1295 |
+
"scored_requests": 10139,
|
| 1296 |
+
"metric": "Brier (lower is better)",
|
| 1297 |
+
"score": 0.1795,
|
| 1298 |
+
"reference_same_cases": null,
|
| 1299 |
+
"median_ms": 73.2,
|
| 1300 |
+
"index_raw": 0.282,
|
| 1301 |
+
"index_skill": 0.282,
|
| 1302 |
+
"coverage": 1.0,
|
| 1303 |
+
"chance": 0.25,
|
| 1304 |
+
"in_index": true
|
| 1305 |
+
},
|
| 1306 |
+
"50": {
|
| 1307 |
+
"catalog_id": 50,
|
| 1308 |
+
"dataset": "Habermas Machine",
|
| 1309 |
+
"requests": 1676,
|
| 1310 |
+
"answered": 1676,
|
| 1311 |
+
"unsupported": 0,
|
| 1312 |
+
"errors": 0,
|
| 1313 |
+
"abstained": 0,
|
| 1314 |
+
"pending": 0,
|
| 1315 |
+
"scored_requests": 1676,
|
| 1316 |
+
"metric": "accuracy",
|
| 1317 |
+
"score": 0.4553,
|
| 1318 |
+
"reference_same_cases": null,
|
| 1319 |
+
"median_ms": 81.4,
|
| 1320 |
+
"index_raw": 0.4553,
|
| 1321 |
+
"index_skill": 0.2094,
|
| 1322 |
+
"coverage": 1.0,
|
| 1323 |
+
"chance": 0.311,
|
| 1324 |
+
"in_index": true
|
| 1325 |
+
},
|
| 1326 |
+
"56": {
|
| 1327 |
+
"catalog_id": 56,
|
| 1328 |
+
"dataset": "PhishNChips phishing decisions",
|
| 1329 |
+
"requests": 2000,
|
| 1330 |
+
"answered": 2000,
|
| 1331 |
+
"unsupported": 0,
|
| 1332 |
+
"errors": 0,
|
| 1333 |
+
"abstained": 0,
|
| 1334 |
+
"pending": 0,
|
| 1335 |
+
"metric": "accuracy",
|
| 1336 |
+
"score": 0.6655,
|
| 1337 |
+
"median_ms": 208.8,
|
| 1338 |
+
"scored_requests": 2000,
|
| 1339 |
+
"index_raw": 0.6655,
|
| 1340 |
+
"index_skill": 0.331,
|
| 1341 |
+
"coverage": 1.0,
|
| 1342 |
+
"chance": 0.5,
|
| 1343 |
+
"in_index": true
|
| 1344 |
+
},
|
| 1345 |
+
"57": {
|
| 1346 |
+
"catalog_id": 57,
|
| 1347 |
+
"dataset": "MMLU-Pro",
|
| 1348 |
+
"requests": 12032,
|
| 1349 |
+
"answered": 12032,
|
| 1350 |
+
"unsupported": 0,
|
| 1351 |
+
"errors": 0,
|
| 1352 |
+
"abstained": 0,
|
| 1353 |
+
"pending": 0,
|
| 1354 |
+
"metric": "accuracy",
|
| 1355 |
+
"score": 0.6137,
|
| 1356 |
+
"median_ms": 30.6,
|
| 1357 |
+
"scored_requests": 12032,
|
| 1358 |
+
"index_raw": 0.6137,
|
| 1359 |
+
"index_skill": 0.5655,
|
| 1360 |
+
"coverage": 1.0,
|
| 1361 |
+
"chance": 0.1109,
|
| 1362 |
+
"in_index": true
|
| 1363 |
+
},
|
| 1364 |
+
"58": {
|
| 1365 |
+
"catalog_id": 58,
|
| 1366 |
+
"dataset": "BBH fixed-option tasks",
|
| 1367 |
+
"requests": 5507,
|
| 1368 |
+
"answered": 5507,
|
| 1369 |
+
"unsupported": 0,
|
| 1370 |
+
"errors": 0,
|
| 1371 |
+
"abstained": 0,
|
| 1372 |
+
"pending": 0,
|
| 1373 |
+
"metric": "accuracy",
|
| 1374 |
+
"score": 0.6784,
|
| 1375 |
+
"median_ms": 23.7,
|
| 1376 |
+
"scored_requests": 5507,
|
| 1377 |
+
"index_raw": 0.6784,
|
| 1378 |
+
"index_skill": 0.5338,
|
| 1379 |
+
"coverage": 1.0,
|
| 1380 |
+
"chance": 0.3101,
|
| 1381 |
+
"in_index": true
|
| 1382 |
+
},
|
| 1383 |
+
"59": {
|
| 1384 |
+
"catalog_id": 59,
|
| 1385 |
+
"dataset": "RAGTruth response-level hallucination",
|
| 1386 |
+
"requests": 2700,
|
| 1387 |
+
"answered": 2700,
|
| 1388 |
+
"unsupported": 0,
|
| 1389 |
+
"errors": 0,
|
| 1390 |
+
"abstained": 0,
|
| 1391 |
+
"pending": 0,
|
| 1392 |
+
"metric": "F1 on hallucinated class",
|
| 1393 |
+
"score": 0.6336,
|
| 1394 |
+
"median_ms": 77.1,
|
| 1395 |
+
"scored_requests": 2700,
|
| 1396 |
+
"index_raw": 0.6336,
|
| 1397 |
+
"index_skill": 0.3776,
|
| 1398 |
+
"coverage": 1.0,
|
| 1399 |
+
"chance": 0.4113,
|
| 1400 |
+
"in_index": true
|
| 1401 |
+
},
|
| 1402 |
+
"61": {
|
| 1403 |
+
"catalog_id": 61,
|
| 1404 |
+
"dataset": "HoVer claim verification",
|
| 1405 |
+
"requests": 4000,
|
| 1406 |
+
"answered": 4000,
|
| 1407 |
+
"unsupported": 0,
|
| 1408 |
+
"errors": 0,
|
| 1409 |
+
"abstained": 0,
|
| 1410 |
+
"pending": 0,
|
| 1411 |
+
"metric": "accuracy",
|
| 1412 |
+
"score": 0.6565,
|
| 1413 |
+
"median_ms": 45.6,
|
| 1414 |
+
"scored_requests": 4000,
|
| 1415 |
+
"index_raw": 0.6565,
|
| 1416 |
+
"index_skill": 0.313,
|
| 1417 |
+
"coverage": 1.0,
|
| 1418 |
+
"chance": 0.5,
|
| 1419 |
+
"in_index": true
|
| 1420 |
+
},
|
| 1421 |
+
"62": {
|
| 1422 |
+
"catalog_id": 62,
|
| 1423 |
+
"dataset": "When2Call MCQ",
|
| 1424 |
+
"requests": 3652,
|
| 1425 |
+
"answered": 3652,
|
| 1426 |
+
"unsupported": 0,
|
| 1427 |
+
"errors": 0,
|
| 1428 |
+
"abstained": 0,
|
| 1429 |
+
"pending": 0,
|
| 1430 |
+
"metric": "accuracy",
|
| 1431 |
+
"score": 0.6451,
|
| 1432 |
+
"median_ms": 79.8,
|
| 1433 |
+
"scored_requests": 3652,
|
| 1434 |
+
"index_raw": 0.6451,
|
| 1435 |
+
"index_skill": 0.5268,
|
| 1436 |
+
"coverage": 1.0,
|
| 1437 |
+
"chance": 0.25,
|
| 1438 |
+
"in_index": true
|
| 1439 |
+
},
|
| 1440 |
+
"64": {
|
| 1441 |
+
"catalog_id": 64,
|
| 1442 |
+
"dataset": "New Yorker caption matching",
|
| 1443 |
+
"requests": 528,
|
| 1444 |
+
"answered": 528,
|
| 1445 |
+
"unsupported": 0,
|
| 1446 |
+
"errors": 0,
|
| 1447 |
+
"abstained": 0,
|
| 1448 |
+
"pending": 0,
|
| 1449 |
+
"metric": "accuracy",
|
| 1450 |
+
"score": 0.625,
|
| 1451 |
+
"median_ms": 25.2,
|
| 1452 |
+
"scored_requests": 528,
|
| 1453 |
+
"index_raw": 0.625,
|
| 1454 |
+
"index_skill": 0.5312,
|
| 1455 |
+
"coverage": 1.0,
|
| 1456 |
+
"chance": 0.2,
|
| 1457 |
+
"in_index": true
|
| 1458 |
+
}
|
| 1459 |
+
},
|
| 1460 |
+
"panel_id": "decision-index-0.2",
|
| 1461 |
+
"note": "Decision Index 0.2 averages 40 benchmarks in five equal-weight areas: each area is the plain mean of its benchmarks and the index is 100 x the mean of the five areas. Each benchmark is chance-corrected first, (score - chance) / (1 - chance) clipped to 0-1, so 0 means random guessing and 100 means perfect. Every score is coverage-adjusted, so an unanswered or unsupported request counts as wrong. ForecastBench enters against its baseline: clip((0.25 - Brier) / 0.25) x coverage, so always predicting 0.5 scores zero. MMLU, ARC-Easy, ARC-Challenge, SimpleBench stay on the board as non-index benchmarks. The six interactive environments are still unrun and stay out. Every entrant on the board has results on all 40 index benchmarks. Point estimates only, no uncertainty intervals yet.",
|
| 1462 |
+
"local_run": {
|
| 1463 |
+
"kit_commit": "19ad28ec9485493cc4f7fc07d91c178f948e6434",
|
| 1464 |
+
"evaluated": "2026-09-25",
|
| 1465 |
+
"note": "Scored locally with the official kit; not a leaderboard submission or result. Requests shared with 0.1 reuse this model's 0.1 predictions; the added requests ran with the same frozen evaluation setup at temperature 1.0. No leaderboard-style exposure penalty is applied: known training exposure (see the model card) stays in these scores.",
|
| 1466 |
+
"without_mmlu_pro": 42.84,
|
| 1467 |
+
"screening": [
|
| 1468 |
+
{
|
| 1469 |
+
"stage": "MiMo",
|
| 1470 |
+
"rows": 123195,
|
| 1471 |
+
"matched": 180,
|
| 1472 |
+
"by_source": {
|
| 1473 |
+
"MMLU-Pro": 0,
|
| 1474 |
+
"SuperGPQA": 176,
|
| 1475 |
+
"BoolQ": 3,
|
| 1476 |
+
"MedMCQA": 1
|
| 1477 |
+
}
|
| 1478 |
+
}
|
| 1479 |
+
],
|
| 1480 |
+
"timing_note": "latency_ms and every benchmark's median_ms are each request's share of batched inference time, allocated by prompt tokens. They are not serial-request or HTTP-serving latency and shouldn't be used for serving-latency comparisons."
|
| 1481 |
+
}
|
| 1482 |
+
}
|