Image-Text-to-Text
Transformers
Safetensors
English
qwen3_5
decision-model
typed-decisions
one-pass
option-probabilities
conversational
Instructions to use thegovind/blink-mimo-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use thegovind/blink-mimo-9b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="thegovind/blink-mimo-9b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("thegovind/blink-mimo-9b") model = AutoModelForMultimodalLM.from_pretrained("thegovind/blink-mimo-9b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use thegovind/blink-mimo-9b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "thegovind/blink-mimo-9b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thegovind/blink-mimo-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/thegovind/blink-mimo-9b
- SGLang
How to use thegovind/blink-mimo-9b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "thegovind/blink-mimo-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thegovind/blink-mimo-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "thegovind/blink-mimo-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thegovind/blink-mimo-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use thegovind/blink-mimo-9b with Docker Model Runner:
docker model run hf.co/thegovind/blink-mimo-9b
Commit ·
7c0a3b8
0
Parent(s):
blink v1.0
Browse files- .dockerignore +2 -0
- .gitattributes +36 -0
- Dockerfile +11 -0
- LICENSE-MiMo.md +23 -0
- LICENSE-Qwen +202 -0
- LICENSE.md +19 -0
- README.md +293 -0
- blink.py +623 -0
- chat_template.jinja +97 -0
- config.json +107 -0
- eval/decision-index-0.1-full.json +1080 -0
- eval/decision-index-0.1-minus-dis.json +1070 -0
- eval/jevbench-public-proxy.json +29 -0
- generation_config.json +10 -0
- merges.txt +0 -0
- model-00001-of-00004.safetensors +3 -0
- model-00002-of-00004.safetensors +3 -0
- model-00003-of-00004.safetensors +3 -0
- model-00004-of-00004.safetensors +3 -0
- model.safetensors.index.json +767 -0
- preprocessor_config.json +26 -0
- processor_config.json +60 -0
- serve.py +167 -0
- tokenizer.json +3 -0
- tokenizer_config.json +33 -0
- video_preprocessor_config.json +21 -0
- vocab.json +0 -0
- weights.sha256 +15 -0
.dockerignore
ADDED
|
@@ -0,0 +1,2 @@
|
|
|
|
|
|
|
|
|
|
| 1 |
+
.git
|
| 2 |
+
.cache
|
.gitattributes
ADDED
|
@@ -0,0 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
*.7z filter=lfs diff=lfs merge=lfs -text
|
| 2 |
+
*.arrow filter=lfs diff=lfs merge=lfs -text
|
| 3 |
+
*.bin filter=lfs diff=lfs merge=lfs -text
|
| 4 |
+
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
| 5 |
+
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
| 6 |
+
*.ftz filter=lfs diff=lfs merge=lfs -text
|
| 7 |
+
*.gz filter=lfs diff=lfs merge=lfs -text
|
| 8 |
+
*.h5 filter=lfs diff=lfs merge=lfs -text
|
| 9 |
+
*.joblib filter=lfs diff=lfs merge=lfs -text
|
| 10 |
+
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
| 11 |
+
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
| 12 |
+
*.model filter=lfs diff=lfs merge=lfs -text
|
| 13 |
+
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
| 14 |
+
*.npy filter=lfs diff=lfs merge=lfs -text
|
| 15 |
+
*.npz filter=lfs diff=lfs merge=lfs -text
|
| 16 |
+
*.onnx filter=lfs diff=lfs merge=lfs -text
|
| 17 |
+
*.ot filter=lfs diff=lfs merge=lfs -text
|
| 18 |
+
*.parquet filter=lfs diff=lfs merge=lfs -text
|
| 19 |
+
*.pb filter=lfs diff=lfs merge=lfs -text
|
| 20 |
+
*.pickle filter=lfs diff=lfs merge=lfs -text
|
| 21 |
+
*.pkl filter=lfs diff=lfs merge=lfs -text
|
| 22 |
+
*.pt filter=lfs diff=lfs merge=lfs -text
|
| 23 |
+
*.pth filter=lfs diff=lfs merge=lfs -text
|
| 24 |
+
*.rar filter=lfs diff=lfs merge=lfs -text
|
| 25 |
+
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
| 26 |
+
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
| 27 |
+
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
| 28 |
+
*.tar filter=lfs diff=lfs merge=lfs -text
|
| 29 |
+
*.tflite filter=lfs diff=lfs merge=lfs -text
|
| 30 |
+
*.tgz filter=lfs diff=lfs merge=lfs -text
|
| 31 |
+
*.wasm filter=lfs diff=lfs merge=lfs -text
|
| 32 |
+
*.xz filter=lfs diff=lfs merge=lfs -text
|
| 33 |
+
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
+
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
+
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
Dockerfile
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# blink server image: this repository's weights and runtime; Hugging Face libraries run in offline mode.
|
| 2 |
+
# hf download thegovind/<model> --revision <sha> --local-dir blink && cd blink
|
| 3 |
+
# docker build -t blink . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink
|
| 4 |
+
# curl -s http://127.0.0.1:8000/healthz # weights_verified, warmup.repeat_identical, kernels
|
| 5 |
+
FROM pytorch/pytorch:2.13.0-cuda12.6-cudnn9-runtime
|
| 6 |
+
ENV PIP_BREAK_SYSTEM_PACKAGES=1 PIP_NO_CACHE_DIR=1 PIP_DISABLE_PIP_VERSION_CHECK=1
|
| 7 |
+
RUN pip install "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.0" "safetensors>=0.4" "huggingface_hub>=1.0"
|
| 8 |
+
COPY . /blink
|
| 9 |
+
ENV HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 HF_HUB_DISABLE_TELEMETRY=1
|
| 10 |
+
EXPOSE 8000
|
| 11 |
+
CMD ["python", "/blink/serve.py", "--model", "/blink", "--host", "0.0.0.0", "--port", "8000"]
|
LICENSE-MiMo.md
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
|
| 2 |
+
|
| 3 |
+
blink-mimo-9b is fine-tuned from [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B)
|
| 4 |
+
(revision `2367e865d009c13ac81713a2878291d33ab28177`). Its model card declares `license: mit` and names
|
| 5 |
+
[Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) as its base model (Apache-2.0; see `LICENSE-Qwen`). The upstream
|
| 6 |
+
repository ships no separate licence file or copyright line; the MIT terms it declares are reproduced below.
|
| 7 |
+
|
| 8 |
+
---
|
| 9 |
+
|
| 10 |
+
MIT License
|
| 11 |
+
|
| 12 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated
|
| 13 |
+
documentation files (the "Software"), to deal in the Software without restriction, including without limitation the
|
| 14 |
+
rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit
|
| 15 |
+
persons to whom the Software is furnished to do so, subject to the following conditions:
|
| 16 |
+
|
| 17 |
+
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the
|
| 18 |
+
Software.
|
| 19 |
+
|
| 20 |
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE
|
| 21 |
+
WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR
|
| 22 |
+
COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR
|
| 23 |
+
OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
LICENSE-Qwen
ADDED
|
@@ -0,0 +1,202 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
|
| 2 |
+
Apache License
|
| 3 |
+
Version 2.0, January 2004
|
| 4 |
+
http://www.apache.org/licenses/
|
| 5 |
+
|
| 6 |
+
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
| 7 |
+
|
| 8 |
+
1. Definitions.
|
| 9 |
+
|
| 10 |
+
"License" shall mean the terms and conditions for use, reproduction,
|
| 11 |
+
and distribution as defined by Sections 1 through 9 of this document.
|
| 12 |
+
|
| 13 |
+
"Licensor" shall mean the copyright owner or entity authorized by
|
| 14 |
+
the copyright owner that is granting the License.
|
| 15 |
+
|
| 16 |
+
"Legal Entity" shall mean the union of the acting entity and all
|
| 17 |
+
other entities that control, are controlled by, or are under common
|
| 18 |
+
control with that entity. For the purposes of this definition,
|
| 19 |
+
"control" means (i) the power, direct or indirect, to cause the
|
| 20 |
+
direction or management of such entity, whether by contract or
|
| 21 |
+
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
| 22 |
+
outstanding shares, or (iii) beneficial ownership of such entity.
|
| 23 |
+
|
| 24 |
+
"You" (or "Your") shall mean an individual or Legal Entity
|
| 25 |
+
exercising permissions granted by this License.
|
| 26 |
+
|
| 27 |
+
"Source" form shall mean the preferred form for making modifications,
|
| 28 |
+
including but not limited to software source code, documentation
|
| 29 |
+
source, and configuration files.
|
| 30 |
+
|
| 31 |
+
"Object" form shall mean any form resulting from mechanical
|
| 32 |
+
transformation or translation of a Source form, including but
|
| 33 |
+
not limited to compiled object code, generated documentation,
|
| 34 |
+
and conversions to other media types.
|
| 35 |
+
|
| 36 |
+
"Work" shall mean the work of authorship, whether in Source or
|
| 37 |
+
Object form, made available under the License, as indicated by a
|
| 38 |
+
copyright notice that is included in or attached to the work
|
| 39 |
+
(an example is provided in the Appendix below).
|
| 40 |
+
|
| 41 |
+
"Derivative Works" shall mean any work, whether in Source or Object
|
| 42 |
+
form, that is based on (or derived from) the Work and for which the
|
| 43 |
+
editorial revisions, annotations, elaborations, or other modifications
|
| 44 |
+
represent, as a whole, an original work of authorship. For the purposes
|
| 45 |
+
of this License, Derivative Works shall not include works that remain
|
| 46 |
+
separable from, or merely link (or bind by name) to the interfaces of,
|
| 47 |
+
the Work and Derivative Works thereof.
|
| 48 |
+
|
| 49 |
+
"Contribution" shall mean any work of authorship, including
|
| 50 |
+
the original version of the Work and any modifications or additions
|
| 51 |
+
to that Work or Derivative Works thereof, that is intentionally
|
| 52 |
+
submitted to Licensor for inclusion in the Work by the copyright owner
|
| 53 |
+
or by an individual or Legal Entity authorized to submit on behalf of
|
| 54 |
+
the copyright owner. For the purposes of this definition, "submitted"
|
| 55 |
+
means any form of electronic, verbal, or written communication sent
|
| 56 |
+
to the Licensor or its representatives, including but not limited to
|
| 57 |
+
communication on electronic mailing lists, source code control systems,
|
| 58 |
+
and issue tracking systems that are managed by, or on behalf of, the
|
| 59 |
+
Licensor for the purpose of discussing and improving the Work, but
|
| 60 |
+
excluding communication that is conspicuously marked or otherwise
|
| 61 |
+
designated in writing by the copyright owner as "Not a Contribution."
|
| 62 |
+
|
| 63 |
+
"Contributor" shall mean Licensor and any individual or Legal Entity
|
| 64 |
+
on behalf of whom a Contribution has been received by Licensor and
|
| 65 |
+
subsequently incorporated within the Work.
|
| 66 |
+
|
| 67 |
+
2. Grant of Copyright License. Subject to the terms and conditions of
|
| 68 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 69 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 70 |
+
copyright license to reproduce, prepare Derivative Works of,
|
| 71 |
+
publicly display, publicly perform, sublicense, and distribute the
|
| 72 |
+
Work and such Derivative Works in Source or Object form.
|
| 73 |
+
|
| 74 |
+
3. Grant of Patent License. Subject to the terms and conditions of
|
| 75 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 76 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 77 |
+
(except as stated in this section) patent license to make, have made,
|
| 78 |
+
use, offer to sell, sell, import, and otherwise transfer the Work,
|
| 79 |
+
where such license applies only to those patent claims licensable
|
| 80 |
+
by such Contributor that are necessarily infringed by their
|
| 81 |
+
Contribution(s) alone or by combination of their Contribution(s)
|
| 82 |
+
with the Work to which such Contribution(s) was submitted. If You
|
| 83 |
+
institute patent litigation against any entity (including a
|
| 84 |
+
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
| 85 |
+
or a Contribution incorporated within the Work constitutes direct
|
| 86 |
+
or contributory patent infringement, then any patent licenses
|
| 87 |
+
granted to You under this License for that Work shall terminate
|
| 88 |
+
as of the date such litigation is filed.
|
| 89 |
+
|
| 90 |
+
4. Redistribution. You may reproduce and distribute copies of the
|
| 91 |
+
Work or Derivative Works thereof in any medium, with or without
|
| 92 |
+
modifications, and in Source or Object form, provided that You
|
| 93 |
+
meet the following conditions:
|
| 94 |
+
|
| 95 |
+
(a) You must give any other recipients of the Work or
|
| 96 |
+
Derivative Works a copy of this License; and
|
| 97 |
+
|
| 98 |
+
(b) You must cause any modified files to carry prominent notices
|
| 99 |
+
stating that You changed the files; and
|
| 100 |
+
|
| 101 |
+
(c) You must retain, in the Source form of any Derivative Works
|
| 102 |
+
that You distribute, all copyright, patent, trademark, and
|
| 103 |
+
attribution notices from the Source form of the Work,
|
| 104 |
+
excluding those notices that do not pertain to any part of
|
| 105 |
+
the Derivative Works; and
|
| 106 |
+
|
| 107 |
+
(d) If the Work includes a "NOTICE" text file as part of its
|
| 108 |
+
distribution, then any Derivative Works that You distribute must
|
| 109 |
+
include a readable copy of the attribution notices contained
|
| 110 |
+
within such NOTICE file, excluding those notices that do not
|
| 111 |
+
pertain to any part of the Derivative Works, in at least one
|
| 112 |
+
of the following places: within a NOTICE text file distributed
|
| 113 |
+
as part of the Derivative Works; within the Source form or
|
| 114 |
+
documentation, if provided along with the Derivative Works; or,
|
| 115 |
+
within a display generated by the Derivative Works, if and
|
| 116 |
+
wherever such third-party notices normally appear. The contents
|
| 117 |
+
of the NOTICE file are for informational purposes only and
|
| 118 |
+
do not modify the License. You may add Your own attribution
|
| 119 |
+
notices within Derivative Works that You distribute, alongside
|
| 120 |
+
or as an addendum to the NOTICE text from the Work, provided
|
| 121 |
+
that such additional attribution notices cannot be construed
|
| 122 |
+
as modifying the License.
|
| 123 |
+
|
| 124 |
+
You may add Your own copyright statement to Your modifications and
|
| 125 |
+
may provide additional or different license terms and conditions
|
| 126 |
+
for use, reproduction, or distribution of Your modifications, or
|
| 127 |
+
for any such Derivative Works as a whole, provided Your use,
|
| 128 |
+
reproduction, and distribution of the Work otherwise complies with
|
| 129 |
+
the conditions stated in this License.
|
| 130 |
+
|
| 131 |
+
5. Submission of Contributions. Unless You explicitly state otherwise,
|
| 132 |
+
any Contribution intentionally submitted for inclusion in the Work
|
| 133 |
+
by You to the Licensor shall be under the terms and conditions of
|
| 134 |
+
this License, without any additional terms or conditions.
|
| 135 |
+
Notwithstanding the above, nothing herein shall supersede or modify
|
| 136 |
+
the terms of any separate license agreement you may have executed
|
| 137 |
+
with Licensor regarding such Contributions.
|
| 138 |
+
|
| 139 |
+
6. Trademarks. This License does not grant permission to use the trade
|
| 140 |
+
names, trademarks, service marks, or product names of the Licensor,
|
| 141 |
+
except as required for reasonable and customary use in describing the
|
| 142 |
+
origin of the Work and reproducing the content of the NOTICE file.
|
| 143 |
+
|
| 144 |
+
7. Disclaimer of Warranty. Unless required by applicable law or
|
| 145 |
+
agreed to in writing, Licensor provides the Work (and each
|
| 146 |
+
Contributor provides its Contributions) on an "AS IS" BASIS,
|
| 147 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
| 148 |
+
implied, including, without limitation, any warranties or conditions
|
| 149 |
+
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
| 150 |
+
PARTICULAR PURPOSE. You are solely responsible for determining the
|
| 151 |
+
appropriateness of using or redistributing the Work and assume any
|
| 152 |
+
risks associated with Your exercise of permissions under this License.
|
| 153 |
+
|
| 154 |
+
8. Limitation of Liability. In no event and under no legal theory,
|
| 155 |
+
whether in tort (including negligence), contract, or otherwise,
|
| 156 |
+
unless required by applicable law (such as deliberate and grossly
|
| 157 |
+
negligent acts) or agreed to in writing, shall any Contributor be
|
| 158 |
+
liable to You for damages, including any direct, indirect, special,
|
| 159 |
+
incidental, or consequential damages of any character arising as a
|
| 160 |
+
result of this License or out of the use or inability to use the
|
| 161 |
+
Work (including but not limited to damages for loss of goodwill,
|
| 162 |
+
work stoppage, computer failure or malfunction, or any and all
|
| 163 |
+
other commercial damages or losses), even if such Contributor
|
| 164 |
+
has been advised of the possibility of such damages.
|
| 165 |
+
|
| 166 |
+
9. Accepting Warranty or Additional Liability. While redistributing
|
| 167 |
+
the Work or Derivative Works thereof, You may choose to offer,
|
| 168 |
+
and charge a fee for, acceptance of support, warranty, indemnity,
|
| 169 |
+
or other liability obligations and/or rights consistent with this
|
| 170 |
+
License. However, in accepting such obligations, You may act only
|
| 171 |
+
on Your own behalf and on Your sole responsibility, not on behalf
|
| 172 |
+
of any other Contributor, and only if You agree to indemnify,
|
| 173 |
+
defend, and hold each Contributor harmless for any liability
|
| 174 |
+
incurred by, or claims asserted against, such Contributor by reason
|
| 175 |
+
of your accepting any such warranty or additional liability.
|
| 176 |
+
|
| 177 |
+
END OF TERMS AND CONDITIONS
|
| 178 |
+
|
| 179 |
+
APPENDIX: How to apply the Apache License to your work.
|
| 180 |
+
|
| 181 |
+
To apply the Apache License to your work, attach the following
|
| 182 |
+
boilerplate notice, with the fields enclosed by brackets "[]"
|
| 183 |
+
replaced with your own identifying information. (Don't include
|
| 184 |
+
the brackets!) The text should be enclosed in the appropriate
|
| 185 |
+
comment syntax for the file format. We also recommend that a
|
| 186 |
+
file or class name and description of purpose be included on the
|
| 187 |
+
same "printed page" as the copyright notice for easier
|
| 188 |
+
identification within third-party archives.
|
| 189 |
+
|
| 190 |
+
Copyright 2026 Alibaba Cloud
|
| 191 |
+
|
| 192 |
+
Licensed under the Apache License, Version 2.0 (the "License");
|
| 193 |
+
you may not use this file except in compliance with the License.
|
| 194 |
+
You may obtain a copy of the License at
|
| 195 |
+
|
| 196 |
+
http://www.apache.org/licenses/LICENSE-2.0
|
| 197 |
+
|
| 198 |
+
Unless required by applicable law or agreed to in writing, software
|
| 199 |
+
distributed under the License is distributed on an "AS IS" BASIS,
|
| 200 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
| 201 |
+
See the License for the specific language governing permissions and
|
| 202 |
+
limitations under the License.
|
LICENSE.md
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
blink model weights: non-commercial research licence
|
| 2 |
+
|
| 3 |
+
These weights are a modification of Qwen/Qwen3.5-9B (Apache License 2.0, Copyright 2026 Alibaba Cloud;
|
| 4 |
+
the full licence text is in LICENSE-Qwen). The modification (fine-tuning with LoRA adapters merged into
|
| 5 |
+
the weights) was made by thegovind.
|
| 6 |
+
|
| 7 |
+
The fine-tuning data included third-party datasets released under different terms, among them
|
| 8 |
+
non-commercial (CC BY-NC 3.0 / 4.0) and share-alike (CC BY-SA 3.0 / 4.0) licences, and sources that
|
| 9 |
+
state no licence (listed in README.md). Whether and how those terms apply to trained weights is unsettled.
|
| 10 |
+
|
| 11 |
+
The author of the modification permits you to use, copy and run it for non-commercial research and
|
| 12 |
+
evaluation only, provided this notice and LICENSE-Qwen are kept with every copy. No licence for commercial
|
| 13 |
+
use is granted, and nothing here grants rights in any third-party data. You remain responsible for
|
| 14 |
+
complying with the Apache License 2.0 for the Qwen base weights and with the terms of the upstream datasets.
|
| 15 |
+
|
| 16 |
+
THE WEIGHTS ARE PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND.
|
| 17 |
+
|
| 18 |
+
This is a personal research release. It is not an official product of any company, and it is not
|
| 19 |
+
affiliated with TypeSafe AI, Alibaba Cloud or the Qwen team.
|
README.md
ADDED
|
@@ -0,0 +1,293 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: blink-research
|
| 4 |
+
license_link: LICENSE.md
|
| 5 |
+
base_model: XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
|
| 6 |
+
library_name: transformers
|
| 7 |
+
pipeline_tag: image-text-to-text
|
| 8 |
+
inference: false
|
| 9 |
+
tags:
|
| 10 |
+
- decision-model
|
| 11 |
+
- typed-decisions
|
| 12 |
+
- one-pass
|
| 13 |
+
- option-probabilities
|
| 14 |
+
language:
|
| 15 |
+
- en
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
# blink-mimo-9b
|
| 19 |
+
|
| 20 |
+
Send a text or JSON `state` and your questions: `choice` picks from up to 255 options, `noul` is yes/no,
|
| 21 |
+
and `score` takes 2–10 ordered levels. Each question gets probabilities over its offered options from
|
| 22 |
+
one forward pass, with no generated text. Long or large multi-question requests may use several batches.
|
| 23 |
+
|
| 24 |
+
**Try it:** [Space demo](https://huggingface.co/spaces/thegovind/blink) — blink-4b and blink-mimo-9b run live ·
|
| 25 |
+
[blink-4b](https://huggingface.co/thegovind/blink-4b) · [blink-27b](https://huggingface.co/thegovind/blink-27b) ·
|
| 26 |
+
[blink-mimo-9b](https://huggingface.co/thegovind/blink-mimo-9b).
|
| 27 |
+
|
| 28 |
+
*Personal research release by thegovind, not an official product of any company. No affiliation with TypeSafe AI,
|
| 29 |
+
Xiaomi, Alibaba Cloud or the Qwen team. Weights are for non-commercial research; see [Licence](#licence).*
|
| 30 |
+
|
| 31 |
+
## Results
|
| 32 |
+
|
| 33 |
+
### Decision Index 0.1 (archived edition)
|
| 34 |
+
|
| 35 |
+
**Full suite: blink-mimo-9b 56.53 vs Jev 1.13.0 59.51.**
|
| 36 |
+
|
| 37 |
+
**Training overlap:** Public train splits included ContractNLI, iSarcasmEval and VAST (the whole Language area), plus Amazon ESCI and Humicroedit. With Language set to Jev's score, this model's index would be 54.96 vs Jev's 59.51. That's arithmetic, not an ablation or a like-for-like comparison with models trained only on synthetic data.
|
| 38 |
+
|
| 39 |
+
| Model | Size class | Decision Index 0.1 | Skill | Breadth |
|
| 40 |
+
|---|---|---:|---:|---:|
|
| 41 |
+
| **blink-mimo-9b** (this model) | 9B | 56.53 | 42.17 | 40.37 |
|
| 42 |
+
| Jev 1.13.0 | closed | 59.51 | 46.26 | 44.79 |
|
| 43 |
+
| Jevfire | 27B | 55.74 | 40.86 | 39.45 |
|
| 44 |
+
| JoshuaSP diffusiongemma (open-jev) | 26B-A4B | 55.56 | 40.84 | 39.19 |
|
| 45 |
+
| Decider 35B-A3B | 35B-A3B | 54.34 | 39.37 | 37.99 |
|
| 46 |
+
| Kev 9B | 9B | 50.48 | 32.96 | 30.54 |
|
| 47 |
+
| Kev 4B | 4B | 47.43 | 28.86 | 25.67 |
|
| 48 |
+
|
| 49 |
+
We ran the complete archived 0.1 suite: 132,422 requests across 37 benchmarks. The headline index averages 19 panel benchmarks. Comparison rows use the 2026-09-22 leaderboard snapshot. We ran the official kit's scorer locally; these aren't leaderboard submissions. The live [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index) moved to 0.2 on 2026-09-24, but the public kit can't build 0.2 yet.
|
| 50 |
+
|
| 51 |
+
No Decision Index 0.2 result is reported for these models. Comparable shared-benchmark results require matched request subsets and the 0.2 metric transformations.
|
| 52 |
+
|
| 53 |
+
Third in our local archived 0.1 comparison, behind blink-27b and Jev and ahead of every open entry in the September 22 snapshot (best: Jevfire).
|
| 54 |
+
|
| 55 |
+
| Area | blink-mimo-9b | Jev 1.13.0 |
|
| 56 |
+
|---|---:|---:|
|
| 57 |
+
| Knowledge & Reasoning | 55.1 | 68.8 |
|
| 58 |
+
| Language Understanding | 70.1 | 62.3 |
|
| 59 |
+
| Retrieval & Classification | 34.8 | 37.0 |
|
| 60 |
+
| Tools & Automation | 70.5 | 73.6 |
|
| 61 |
+
| Arts & Human Judgment | 52.1 | 56.2 |
|
| 62 |
+
|
| 63 |
+
<details><summary>Per benchmark (19 panel benchmarks, 0.1)</summary>
|
| 64 |
+
|
| 65 |
+
| Area | Benchmark | This model | Jev 1.13.0 |
|
| 66 |
+
|---|---|---:|---:|
|
| 67 |
+
| Knowledge | MMLU | 0.802 | 0.917 |
|
| 68 |
+
| Knowledge | GPQA Diamond | 0.408 | 0.783 |
|
| 69 |
+
| Knowledge | GSM8K | 0.658 | 0.799 |
|
| 70 |
+
| Knowledge | CRUXEval | 0.547 | 0.730 |
|
| 71 |
+
| Knowledge | CLadder | 0.661 | 0.726 |
|
| 72 |
+
| Knowledge | ChessBench | 0.229 | 0.172 |
|
| 73 |
+
| Language | ContractNLI | 0.817 | 0.717 |
|
| 74 |
+
| Language | iSarcasmEval | 0.506 | 0.505 |
|
| 75 |
+
| Language | VAST | 0.780 | 0.646 |
|
| 76 |
+
| Retrieval | BRIGHT | 0.177 | 0.187 |
|
| 77 |
+
| Retrieval | Amazon ESCI | 0.520 | 0.552 |
|
| 78 |
+
| Tools | BFCL | 0.893 | 0.958 |
|
| 79 |
+
| Tools | ToolRet | 0.422 | 0.450 |
|
| 80 |
+
| Tools | RouterBench | 0.799 | 0.799 |
|
| 81 |
+
| Arts | BPoMP | 0.841 | 0.906 |
|
| 82 |
+
| Arts | Humicroedit | 0.638 | 0.619 |
|
| 83 |
+
| Arts | POP909-CL | 0.076 | 0.181 |
|
| 84 |
+
| Arts | cfcolor | 0.597 | 0.647 |
|
| 85 |
+
| Arts | Habermas Machine | 0.455 | 0.459 |
|
| 86 |
+
|
| 87 |
+
</details>
|
| 88 |
+
|
| 89 |
+
### JevBench: public items only
|
| 90 |
+
|
| 91 |
+
| Model | Easy | Standard | Hard | Hard ECE | Official score |
|
| 92 |
+
|---|---:|---:|---:|---:|---|
|
| 93 |
+
| **blink-mimo-9b** | 48/48 | 70/72 | 77/111 (0.694) | 0.136 | not submitted |
|
| 94 |
+
|
| 95 |
+
Same public-item harness as the other models. These are development results, not official scores or predictions; no official rank or parity is claimed. blink-4b is the JevBench entry. MiMo's hard ECE is 0.136 vs blink-4b's 0.067; it's less calibrated here. Its weakest public hard cases are date/number reasoning, trade-offs and long policies.
|
| 96 |
+
|
| 97 |
+
### Picking the next click (our probe)
|
| 98 |
+
|
| 99 |
+
| Input | blink-mimo-9b | MiMo base | blink-4b |
|
| 100 |
+
|---|---:|---:|---:|
|
| 101 |
+
| Page text | 53.8% | 49.0% | 55.4% |
|
| 102 |
+
| Screenshot | 48.4% | 43.0% | — |
|
| 103 |
+
|
| 104 |
+
Our harness and sampling on 500 Multimodal-Mind2Web test steps with five offered elements (chance 20%), not the official evaluation. Five choices are easier than ranking a whole page; blink never trained on web actions. Training lifted MiMo by about five points in both modes, but text still beats screenshots; blink-4b is as good with page text.
|
| 105 |
+
|
| 106 |
+
## What we changed in the network
|
| 107 |
+
|
| 108 |
+
| | What ships |
|
| 109 |
+
|---|---|
|
| 110 |
+
| Backbone | MiMo-V2.6-Distill-Qwen-9B (upstream revision 2367e86), a Qwen3.5-9B fine-tune; 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 4096, untied embeddings |
|
| 111 |
+
| Vision | 27 encoder blocks, hidden 1152 projected to 4096; all 333 vision tensors unchanged |
|
| 112 |
+
| Tuned | 43.3M LoRA parameters; all 248 targeted tensors changed, all 179 other language-model tensors bit-identical to MiMo |
|
| 113 |
+
| Final weights | Merged and grafted into MiMo; vision-language reload with no missing or unexpected keys |
|
| 114 |
+
| Serving | 9.41B parameters in the full checkpoint (0.46B vision); 8.95B text parameters serving decisions (17.9 GB in bf16); 18.8 GB total |
|
| 115 |
+
|
| 116 |
+
**LoRA targets (rank 16, alpha 32, every language-model layer):** full-attention `q_proj`, `k_proj`, `v_proj`, `o_proj`;
|
| 117 |
+
Gated DeltaNet `in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `out_proj`; and every MLP's
|
| 118 |
+
`gate_proj`, `up_proj`, `down_proj`. Token embeddings, all norms and `lm_head` stayed frozen.
|
| 119 |
+
The trained adapters were merged into the text weights.
|
| 120 |
+
|
| 121 |
+
MiMo's vision tower stays byte-for-byte unchanged. The full checkpoint loads with MiMo's processor and transformers' image-text-to-text classes; `blink.py` and `serve.py` use its text side only. Screenshot results above come from our separate probe harness, not `blink.py`.
|
| 122 |
+
|
| 123 |
+
**The real cut is at readout:** no text generation. One prompt pass; next-token logits from only the
|
| 124 |
+
offered option-label rows of `lm_head` (verified single tokens A–Z, then two-letter labels), computed in
|
| 125 |
+
FP32 and softmaxed over those letters. The rest of the vocabulary is ignored.
|
| 126 |
+
|
| 127 |
+
**Objective:** "calibration-oriented decision post-training" is plain supervised fine-tuning.
|
| 128 |
+
Cross-entropy uses each row's target distribution: code-computed exact probabilities, probability
|
| 129 |
+
targets in teacher-written questions kept after a blind re-solve by that same teacher agreed, and
|
| 130 |
+
one-hot labels otherwise. Choice and yes/no options and letter assignments are
|
| 131 |
+
reshuffled each epoch; score levels keep their order. Jev's RLCD recipe isn't public; we didn't
|
| 132 |
+
use or reproduce it. No RL or preference optimisation.
|
| 133 |
+
|
| 134 |
+
## The climb
|
| 135 |
+
|
| 136 |
+
One pre-registered run, one epoch; no other MiMo variant was trained or picked. The gate was DI-S ≥ 55 before training: pass it, then read the full suite and release. DI-S intervals are a few points wide, so this is a narrow gate pass, not a claim of superiority.
|
| 137 |
+
|
| 138 |
+
| Step | DI-S | Full 0.1 | Outside DI-S | Why |
|
| 139 |
+
|---|---:|---:|---:|---|
|
| 140 |
+
| MiMo base, zero-shot | 47.09 | — | — | Baseline before decision training. |
|
| 141 |
+
| **blink-mimo-9b** (123,195 question rows; lr 5e-5; 615 steps) | 55.3 | 56.53 | 56.60 | Program-labelled reasoning, teacher questions and judge data cleared the gate; shipped. |
|
| 142 |
+
|
| 143 |
+
DI-S areas: Knowledge 55.6, Language 66.4, Retrieval 33.0, Tools 70.7, Arts 50.8.
|
| 144 |
+
|
| 145 |
+
## Use
|
| 146 |
+
|
| 147 |
+
```python
|
| 148 |
+
# pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
|
| 149 |
+
import os, sys
|
| 150 |
+
from huggingface_hub import hf_hub_download
|
| 151 |
+
|
| 152 |
+
os.environ["BLINK_MODEL"] = "thegovind/blink-mimo-9b"
|
| 153 |
+
os.environ["BLINK_REVISION"] = "v1.0"
|
| 154 |
+
sys.path.insert(0, os.path.dirname(hf_hub_download("thegovind/blink-mimo-9b", "blink.py", revision="v1.0")))
|
| 155 |
+
import blink
|
| 156 |
+
|
| 157 |
+
out = blink.decide(
|
| 158 |
+
"Order #4411 arrived with a cracked screen. I want my money back, not another one.",
|
| 159 |
+
{
|
| 160 |
+
"intent": {
|
| 161 |
+
"type": "choice",
|
| 162 |
+
"instructions": "What does the customer want?",
|
| 163 |
+
"criteria": {"refund": "Money back", "replacement": "A new unit", "info": "Information only"},
|
| 164 |
+
},
|
| 165 |
+
"urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
|
| 166 |
+
"anger": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]},
|
| 167 |
+
},
|
| 168 |
+
)
|
| 169 |
+
print(out["answers"]["intent"]["probabilities"])
|
| 170 |
+
```
|
| 171 |
+
|
| 172 |
+
## Run it as a server
|
| 173 |
+
|
| 174 |
+
`serve.py` accepts Jev-compatible `POST /v1/systemone` (`{state, questions}` → `{answers, usage}`);
|
| 175 |
+
JevBench's stock `typesafe` adapter and the Decision Index kit's `http` engine use this format.
|
| 176 |
+
`GET /healthz` reports startup checks.
|
| 177 |
+
|
| 178 |
+
```sh
|
| 179 |
+
pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
|
| 180 |
+
hf download thegovind/blink-mimo-9b --revision v1.0 --local-dir blink-mimo-9b
|
| 181 |
+
python blink-mimo-9b/serve.py --model ./blink-mimo-9b --port 8000
|
| 182 |
+
```
|
| 183 |
+
|
| 184 |
+
In another terminal: `curl -s http://127.0.0.1:8000/healthz`.
|
| 185 |
+
|
| 186 |
+
**Or use Docker** from the downloaded folder:
|
| 187 |
+
|
| 188 |
+
```sh
|
| 189 |
+
cd blink-mimo-9b
|
| 190 |
+
docker build -t blink-mimo-9b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-mimo-9b
|
| 191 |
+
```
|
| 192 |
+
|
| 193 |
+
<details><summary>Health, limits and weights</summary>
|
| 194 |
+
|
| 195 |
+
- `/healthz` reports `weights_verified` (weight, config and tokenizer files listed in `weights.sha256`
|
| 196 |
+
are hashed before serving; a mismatch stops startup), `warmup.repeat_identical` (two matching warm-up
|
| 197 |
+
answers), `kernels` (fast path or slower fallback without flash-linear-attention), `versions` and `hub_offline`.
|
| 198 |
+
- Limits: 255 options per choice, 2–10 score levels, 131,072 input tokens per question and 512 questions
|
| 199 |
+
per request. Over-limit requests get HTTP 422 with the reason; nothing is truncated.
|
| 200 |
+
- Requests run one at a time. Questions are batched; each batch takes one forward pass (large requests
|
| 201 |
+
can take more than one). Serving the downloaded folder or Docker image enables Hugging Face offline
|
| 202 |
+
mode before model loading (`hub_offline: true`). The server doesn't otherwise restrict network access.
|
| 203 |
+
- blink-mimo-9b weights are 18.8 GB in bf16. Long prompts need more memory.
|
| 204 |
+
|
| 205 |
+
</details>
|
| 206 |
+
|
| 207 |
+
<details><summary>Model and probability readout</summary>
|
| 208 |
+
|
| 209 |
+
| | |
|
| 210 |
+
|---|---|
|
| 211 |
+
| Base | [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) (MIT declared in the model card; see LICENSE-MiMo.md), full vision-language weights |
|
| 212 |
+
| Adaptation | LoRA r16 / alpha 32 across the language model's attention, Gated DeltaNet and MLP; merged into the full MiMo checkpoint |
|
| 213 |
+
| Readout | Verified single-token labels (A–Z, then two-letter labels); FP32 softmax of next-token logits / T over the offered labels |
|
| 214 |
+
| Temperature | 1.0 (not fitted) |
|
| 215 |
+
| Input limit | 131,072 tokens per question; longest evaluated prompt: 37,906 tokens; longer inputs are refused, never truncated |
|
| 216 |
+
| Runtime | `blink.py` builds prompts and reads out probabilities for text decisions |
|
| 217 |
+
|
| 218 |
+
These are option-conditional model probabilities, not certified chances of being right. Calibration can shift
|
| 219 |
+
across tasks, domains and option sets. For `choice`, `confidence = (p_max − 1/K)/(1 − 1/K)` measures
|
| 220 |
+
concentration, not correctness. For `score`, `score` is the expected 0-based level and `choice` is the
|
| 221 |
+
most likely level.
|
| 222 |
+
|
| 223 |
+
</details>
|
| 224 |
+
|
| 225 |
+
<details><summary>Training and data</summary>
|
| 226 |
+
|
| 227 |
+
### How it was trained
|
| 228 |
+
|
| 229 |
+
| Stage | Question rows | Mix |
|
| 230 |
+
|---|---:|---|
|
| 231 |
+
| MiMo | 123,195 | 61,394 public-source · 23,894 program-generated reasoning · 12,000 decision worlds · 7,860 teacher-written question rows · 7,000 judge-style · 6,000 chess move choices · 5,047 exact-probability worlds |
|
| 232 |
+
|
| 233 |
+
One pre-registered supervised run, one epoch (lr 5e-5, 615 steps). The mix joins the 27B's stage-1 public and exact-probability rows with its stage-2 reasoning, chess and decision worlds, plus judge-style data. Qwen3.8-27B wrote the teacher documents and typed questions. Public train splits include ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, ANLI, WANLI, BoolQ, BANKING77, MedMCQA, AQuA, MMLU auxiliary train, SciQ and SuperGPQA.
|
| 234 |
+
|
| 235 |
+
### Data sources and licences
|
| 236 |
+
|
| 237 |
+
| Source | Licence |
|
| 238 |
+
|---|---|
|
| 239 |
+
| MMLU auxiliary train, CommonsenseQA, GSM8K | MIT |
|
| 240 |
+
| AQuA-RAT, Amazon ESCI | Apache-2.0 |
|
| 241 |
+
| searchless_chess | data CC BY 4.0 (Lichess-derived portions CC0); code Apache-2.0 |
|
| 242 |
+
| MedMCQA | Apache-2.0 (dataset card) |
|
| 243 |
+
| SuperGPQA | ODC-BY |
|
| 244 |
+
| WANLI, ContractNLI, BANKING77 | CC BY 4.0 |
|
| 245 |
+
| ARC | CC BY-SA 4.0 |
|
| 246 |
+
| BoolQ, Dolly-15k | CC BY-SA 3.0 |
|
| 247 |
+
| ANLI | CC BY-NC 4.0 |
|
| 248 |
+
| SciQ | CC BY-NC 3.0 |
|
| 249 |
+
| iSarcasmEval | MIT (upstream repository licence) |
|
| 250 |
+
| VAST, Humicroedit, OpenBookQA | None stated by source |
|
| 251 |
+
| Our code-generated worlds and teacher-written documents (Qwen3.8-27B) | See LICENSE.md |
|
| 252 |
+
|
| 253 |
+
These are source-repository licences; they don't settle rights in every underlying text.
|
| 254 |
+
|
| 255 |
+
</details>
|
| 256 |
+
|
| 257 |
+
<details><summary>Evaluation notes and limits</summary>
|
| 258 |
+
|
| 259 |
+
### Evaluation notes
|
| 260 |
+
|
| 261 |
+
- **Scorer parity.** `blink.py` and the lab scorer agree within 1e-7 on this checkpoint (p99 |Δp| 1.4e-8); the published graft matches the evaluated adapter on JevBench's 231 public items (0 argmax changes, max |Δp| 0.0). The kit's per-request timer on 1,000 random suite requests served one at a time measured 68.3 ms median via HTTP vs 66.9 ms in-process. From a fresh Hub download and install at the pinned revision, `serve.py` with JevBench's stock adapter returned easy 48/48, standard 70/72 and hard 77/111, matching the evaluation's answers (0 argmax changes; max |Δp| about 0.03).
|
| 262 |
+
- **Selection.** The prompt format came from earlier DI-S reads; this MiMo run used the sample as a pre-registered gate. The full suite followed that gate; it scored 56.60 on the 129,422 requests outside DI-S, which were not used for selection.
|
| 263 |
+
- **Training overlap.** Public train splits also used by the 0.1 index: ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, ChessBench (searchless_chess training positions; none of the 5,000 test positions), GSM8K (train split; solution-checking items). We also used ANLI and BANKING77 train splits; they're in the 0.1 suite but outside its index, and both are in the 0.2 panel. No identical suite test row was used; the overlap audit below covers shared passages.
|
| 264 |
+
- **Partitions.** Public-source data included training and development partitions.
|
| 265 |
+
- **Final-mixture audit.** Rechecked every question row (including teacher-written rows) against the complete 0.1 suite (132,422 requests) and JevBench's 231 public items. The checks used normalised text of at least 30 characters and 13-word passages. No public JevBench item matched; no chess position is shared.
|
| 266 |
+
- **Suite overlap.** 16 BANKING77/VAST training rows share a 13-word passage with 31 suite requests: 23 of VAST's 3,006 and 8 of BANKING77's 3,080. Two VAST training posts are near-duplicates of a test post; none of these texts is identical. Dropping those requests leaves the index at 56.53 (VAST 0.7805 → 0.7803); BANKING77 is outside the index.
|
| 267 |
+
- **Audit limits.** Semantic or pretraining overlap can't be ruled out; private JevBench items weren't available to check.
|
| 268 |
+
- **Generated reasoning.** Our programs computed the labels for CRUXEval-style code and CLadder-style causal questions; no items from those benchmarks were used. We didn't reuse the suite's GSM8K distractors.
|
| 269 |
+
- **Teacher documents.** We kept Qwen3.8-27B's documents only if a fresh blind solve by that same teacher agreed with the answer. That's an agreement filter, not independent verification.
|
| 270 |
+
- The repo ships no benchmark items, GPQA text, JevBench items or teacher traces.
|
| 271 |
+
- **No MMLU-Pro or GPQA.** No rows from either were used in this fine-tune; SuperGPQA is a separate source.
|
| 272 |
+
|
| 273 |
+
### Limits
|
| 274 |
+
|
| 275 |
+
- English-centric. Training included Arabic iSarcasmEval rows; on the 0.1 suite, Arabic task A scored 0.321 and task C pairs 0.79. Broader multilingual performance hasn't been established.
|
| 276 |
+
- Doesn't chat or explain answers.
|
| 277 |
+
- Text in the state can sway the answer.
|
| 278 |
+
- Its weakest public hard cases are date/number reasoning, trade-offs and long policies.
|
| 279 |
+
|
| 280 |
+
</details>
|
| 281 |
+
|
| 282 |
+
## Licence
|
| 283 |
+
|
| 284 |
+
The MiMo model card declares MIT without a separate upstream licence file or copyright line (`LICENSE-MiMo.md`); its Qwen/Qwen3.5-9B base is Apache-2.0 (`LICENSE-Qwen`). The blink weights are for **non-commercial research and evaluation
|
| 285 |
+
only** (`LICENSE.md`); commercial use isn't licensed. Training used non-commercial, share-alike and unlicensed
|
| 286 |
+
sources (see the table above). It's unsettled whether their terms reach the weights, so check upstream terms too.
|
| 287 |
+
`blink.py`, `serve.py` and the Dockerfile are Apache-2.0.
|
| 288 |
+
|
| 289 |
+
<details><summary>Credits</summary>
|
| 290 |
+
|
| 291 |
+
Xiaomi MiMo (MiMo base) and the Qwen team (Qwen base). SemIf (MIT) for the evidence/criterion/options prompt layout. The Decision Index kit (MIT) and JevBench (MIT) for evaluation.
|
| 292 |
+
|
| 293 |
+
</details>
|
blink.py
ADDED
|
@@ -0,0 +1,623 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""blink — one-pass typed decisions.
|
| 2 |
+
|
| 3 |
+
A request is a `state` plus a map of typed questions. Every question is rendered
|
| 4 |
+
independently with the fixed "semif" template the model was trained on, the questions run
|
| 5 |
+
as right-padded prefills (packed into batches of at most BLINK_TOKEN_BUDGET padded tokens,
|
| 6 |
+
all inside one GPU call), and the answer is read from the next-token distribution over the
|
| 7 |
+
offered option labels. No tokens are generated.
|
| 8 |
+
|
| 9 |
+
Engines (BLINK_ENGINE): "torch" runs the model; "replay" serves label logits the real
|
| 10 |
+
model produced for the bundled examples (hosting without the model); "hybrid" serves those
|
| 11 |
+
recordings for the bundled examples and runs the model for everything else; "mock" is a
|
| 12 |
+
deterministic stand-in with fabricated numbers (BLINK_MOCK=1 also selects it). By default:
|
| 13 |
+
hybrid when CUDA and a recording are both present, torch when only CUDA is, replay otherwise.
|
| 14 |
+
|
| 15 |
+
BLINK_MODELS="repo[@revision],repo2" serves several models side by side (decide(..., model=...));
|
| 16 |
+
the default model's recording is replay.json, another model's replay-<repo name>.json, and a
|
| 17 |
+
recording is only used by the model that made it.
|
| 18 |
+
"""
|
| 19 |
+
|
| 20 |
+
from __future__ import annotations
|
| 21 |
+
|
| 22 |
+
import hashlib
|
| 23 |
+
import itertools
|
| 24 |
+
import json
|
| 25 |
+
import math
|
| 26 |
+
import os
|
| 27 |
+
import re
|
| 28 |
+
import string
|
| 29 |
+
import time
|
| 30 |
+
|
| 31 |
+
# --- constants the deployment sets -------------------------------------------------
|
| 32 |
+
|
| 33 |
+
MODEL_ID = os.environ.get("BLINK_MODEL") or "thegovind/blink-4b"
|
| 34 |
+
MODEL_ID_27B = "thegovind/blink-27b"
|
| 35 |
+
MODEL_REVISION = os.environ.get("BLINK_REVISION") or None
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def _model_specs() -> list[tuple[str, str | None]]:
|
| 39 |
+
"""BLINK_MODELS serves several models side by side: comma-separated "repo" or "repo@revision",
|
| 40 |
+
the first one the default. Unset, the single BLINK_MODEL / BLINK_REVISION pair is served."""
|
| 41 |
+
specs = []
|
| 42 |
+
for item in os.environ.get("BLINK_MODELS", "").split(","):
|
| 43 |
+
repo, _, rev = item.strip().partition("@")
|
| 44 |
+
if repo.strip():
|
| 45 |
+
specs.append((repo.strip(), rev.strip() or None))
|
| 46 |
+
return specs or [(MODEL_ID, MODEL_REVISION)]
|
| 47 |
+
|
| 48 |
+
|
| 49 |
+
MODEL_SPECS = _model_specs()
|
| 50 |
+
MODEL_ID, MODEL_REVISION = MODEL_SPECS[0]
|
| 51 |
+
|
| 52 |
+
# Default release temperature: 1.0, not fitted; callers may override it.
|
| 53 |
+
TEMPERATURE = float(os.environ.get("BLINK_TEMPERATURE", "1.0"))
|
| 54 |
+
|
| 55 |
+
MAX_OPTIONS = 255
|
| 56 |
+
MAX_QUESTIONS = 512
|
| 57 |
+
MAX_INPUT_TOKENS = int(os.environ.get("BLINK_MAX_INPUT_TOKENS", "131072")) # as evaluated
|
| 58 |
+
TOKEN_BUDGET = int(os.environ.get("BLINK_TOKEN_BUDGET", "32768")) # padded tokens per forward
|
| 59 |
+
GPU_DURATION = int(os.environ.get("BLINK_GPU_DURATION", "60"))
|
| 60 |
+
|
| 61 |
+
SYSTEM = (
|
| 62 |
+
"Apply the supplied criterion to the supplied evidence. Choose exactly one listed option. "
|
| 63 |
+
"Respond with only its uppercase letter, with no explanation or reasoning."
|
| 64 |
+
)
|
| 65 |
+
|
| 66 |
+
LABEL_POOL = list(string.ascii_uppercase) + [
|
| 67 |
+
"".join(p) for p in itertools.product(string.ascii_uppercase, repeat=2)
|
| 68 |
+
]
|
| 69 |
+
GENERIC_KEY = re.compile(r"^(?:[A-Za-z]{1,2}|option[_ ]?\d+|opt\d+|\d+)$")
|
| 70 |
+
|
| 71 |
+
MOCK = os.environ.get("BLINK_MOCK", "") not in ("", "0", "false", "False")
|
| 72 |
+
ENGINE = "mock" if MOCK else os.environ.get("BLINK_ENGINE", "").strip().lower()
|
| 73 |
+
HERE = os.path.dirname(os.path.abspath(__file__))
|
| 74 |
+
REPLAY_PATH = os.environ.get("BLINK_REPLAY", os.path.join(HERE, "replay.json"))
|
| 75 |
+
|
| 76 |
+
|
| 77 |
+
def replay_path(model_id: str | None = None) -> str:
|
| 78 |
+
"""The default model's recording is replay.json (or BLINK_REPLAY); any other served model's
|
| 79 |
+
is replay-<repo name>.json beside this file."""
|
| 80 |
+
if model_id is None or model_id == MODEL_ID:
|
| 81 |
+
return REPLAY_PATH
|
| 82 |
+
return os.path.join(HERE, f"replay-{model_id.rstrip('/').split('/')[-1]}.json")
|
| 83 |
+
|
| 84 |
+
|
| 85 |
+
class BlinkError(ValueError):
|
| 86 |
+
"""A request that cannot be rendered."""
|
| 87 |
+
|
| 88 |
+
|
| 89 |
+
class ReplayMiss(BlinkError):
|
| 90 |
+
"""Replay mode has no recorded output for this request."""
|
| 91 |
+
|
| 92 |
+
|
| 93 |
+
# --- rendering (pure python; the model was trained on exactly this) -----------------
|
| 94 |
+
|
| 95 |
+
|
| 96 |
+
def text(x) -> str:
|
| 97 |
+
if x is None:
|
| 98 |
+
return ""
|
| 99 |
+
if isinstance(x, str):
|
| 100 |
+
return x
|
| 101 |
+
return json.dumps(x, ensure_ascii=False, indent=2)
|
| 102 |
+
|
| 103 |
+
|
| 104 |
+
def option_text(key, desc) -> str:
|
| 105 |
+
k = str(key)
|
| 106 |
+
if desc is None or (isinstance(desc, str) and not desc.strip()):
|
| 107 |
+
return k
|
| 108 |
+
d = text(desc)
|
| 109 |
+
if GENERIC_KEY.match(k) or k.strip().lower() == d.strip().lower():
|
| 110 |
+
return d
|
| 111 |
+
return f"{k}: {d}"
|
| 112 |
+
|
| 113 |
+
|
| 114 |
+
def question_options(q: dict) -> list[tuple[str, str]]:
|
| 115 |
+
"""[(option_key, display_text)] in request order."""
|
| 116 |
+
qtype = q.get("type")
|
| 117 |
+
crit = q.get("criteria")
|
| 118 |
+
if qtype == "choice":
|
| 119 |
+
if isinstance(crit, dict):
|
| 120 |
+
items = [(str(k), option_text(k, v)) for k, v in crit.items()]
|
| 121 |
+
elif isinstance(crit, list):
|
| 122 |
+
items = [(str(k), str(k)) for k in crit]
|
| 123 |
+
else:
|
| 124 |
+
raise BlinkError("choice needs criteria")
|
| 125 |
+
elif qtype == "noul":
|
| 126 |
+
t = f = None
|
| 127 |
+
if isinstance(crit, dict):
|
| 128 |
+
t = crit.get("true", crit.get("yes"))
|
| 129 |
+
f = crit.get("false", crit.get("no"))
|
| 130 |
+
items = [
|
| 131 |
+
("yes", "Yes" + (f" — {text(t)}" if t else "")),
|
| 132 |
+
("no", "No" + (f" — {text(f)}" if f else "")),
|
| 133 |
+
]
|
| 134 |
+
elif qtype == "score":
|
| 135 |
+
if not isinstance(crit, list) or not 2 <= len(crit) <= 10:
|
| 136 |
+
raise BlinkError("a score takes 2 to 10 levels")
|
| 137 |
+
items = [
|
| 138 |
+
(str(i), f"Level {i}: {text(c)}" if c is not None else f"Level {i}")
|
| 139 |
+
for i, c in enumerate(crit)
|
| 140 |
+
]
|
| 141 |
+
else:
|
| 142 |
+
raise BlinkError(f"unsupported question type {qtype!r}")
|
| 143 |
+
keys = [k for k, _ in items]
|
| 144 |
+
if len(set(keys)) != len(keys):
|
| 145 |
+
raise BlinkError("duplicate option keys")
|
| 146 |
+
if not 1 <= len(items) <= MAX_OPTIONS:
|
| 147 |
+
raise BlinkError(f"{len(items)} options per choice; supported 1-{MAX_OPTIONS}")
|
| 148 |
+
return items
|
| 149 |
+
|
| 150 |
+
|
| 151 |
+
def user_message(state, q: dict, labels: list[str], items: list[tuple[str, str]]) -> str:
|
| 152 |
+
return json.dumps(
|
| 153 |
+
{
|
| 154 |
+
"evidence": state,
|
| 155 |
+
"criterion": text(q.get("instructions")).strip(),
|
| 156 |
+
"options": [
|
| 157 |
+
{"letter": lab, "description": d} for lab, (_, d) in zip(labels, items)
|
| 158 |
+
],
|
| 159 |
+
},
|
| 160 |
+
ensure_ascii=False,
|
| 161 |
+
)
|
| 162 |
+
|
| 163 |
+
|
| 164 |
+
def validate(questions) -> None:
|
| 165 |
+
if not isinstance(questions, dict) or not questions:
|
| 166 |
+
raise BlinkError("questions must be a non-empty object")
|
| 167 |
+
if len(questions) > MAX_QUESTIONS:
|
| 168 |
+
raise BlinkError(f"{len(questions)} questions; supported 1-{MAX_QUESTIONS}")
|
| 169 |
+
for qkey, q in questions.items():
|
| 170 |
+
if not isinstance(q, dict):
|
| 171 |
+
raise BlinkError(f"question {qkey!r} must be an object")
|
| 172 |
+
question_options(q)
|
| 173 |
+
|
| 174 |
+
|
| 175 |
+
def as_state(state):
|
| 176 |
+
"""A state that looks like JSON is passed through as JSON; anything else is text."""
|
| 177 |
+
if not isinstance(state, str) or not state.strip().startswith(("{", "[")):
|
| 178 |
+
return state
|
| 179 |
+
try:
|
| 180 |
+
return json.loads(state)
|
| 181 |
+
except json.JSONDecodeError:
|
| 182 |
+
return state
|
| 183 |
+
|
| 184 |
+
|
| 185 |
+
# --- answer assembly ----------------------------------------------------------------
|
| 186 |
+
|
| 187 |
+
|
| 188 |
+
def answer_for(q: dict, keys: list[str], probs: list[float]) -> dict:
|
| 189 |
+
p = {k: v for k, v in zip(keys, probs)}
|
| 190 |
+
z = sum(p.values()) or 1.0
|
| 191 |
+
p = {k: v / z for k, v in p.items()}
|
| 192 |
+
qtype = q["type"]
|
| 193 |
+
if qtype == "choice":
|
| 194 |
+
order = list(p)
|
| 195 |
+
best = max(order, key=lambda k: (p[k], -order.index(k)))
|
| 196 |
+
K = len(order)
|
| 197 |
+
conf = (p[best] - 1 / K) / (1 - 1 / K) if K > 1 else 1.0
|
| 198 |
+
return {"type": "choice", "choice": best, "probabilities": p, "confidence": conf}
|
| 199 |
+
if qtype == "noul":
|
| 200 |
+
return {"type": "noul", "noul": p["yes"], "probabilities": p}
|
| 201 |
+
levels = [str(i) for i in range(len(q["criteria"]))]
|
| 202 |
+
ev = sum(int(k) * p[k] for k in levels)
|
| 203 |
+
return {
|
| 204 |
+
"type": "score",
|
| 205 |
+
"score": ev,
|
| 206 |
+
"probabilities": {k: p[k] for k in levels},
|
| 207 |
+
"legend": {k: text(q["criteria"][int(k)]) for k in levels},
|
| 208 |
+
"choice": max(levels, key=lambda k: (p[k], -int(k))),
|
| 209 |
+
}
|
| 210 |
+
|
| 211 |
+
|
| 212 |
+
def softmax(logits: list[float], temperature: float) -> list[float]:
|
| 213 |
+
z = [v / temperature for v in logits]
|
| 214 |
+
m = max(z)
|
| 215 |
+
e = [math.exp(v - m) for v in z]
|
| 216 |
+
s = sum(e)
|
| 217 |
+
return [v / s for v in e]
|
| 218 |
+
|
| 219 |
+
|
| 220 |
+
# --- mock engine --------------------------------------------------------------------
|
| 221 |
+
|
| 222 |
+
_WORD = re.compile(r"[a-z0-9']+")
|
| 223 |
+
_STOP = frozenset(
|
| 224 |
+
"a an the of to and or is are was were be been it its this that for in on at "
|
| 225 |
+
"with as by from not no yes if then than there here we you they i".split()
|
| 226 |
+
)
|
| 227 |
+
|
| 228 |
+
|
| 229 |
+
def _terms(s: str) -> set[str]:
|
| 230 |
+
return {w for w in _WORD.findall(s.lower()) if w not in _STOP and len(w) > 2}
|
| 231 |
+
|
| 232 |
+
|
| 233 |
+
class MockEngine:
|
| 234 |
+
"""Deterministic stand-in: lexical overlap plus a stable pseudo-random jitter.
|
| 235 |
+
|
| 236 |
+
Values are fabricated. They exist so the interface can be exercised without a GPU.
|
| 237 |
+
A caller may pass `bias` to shape a bundled demo; the real engine ignores it.
|
| 238 |
+
"""
|
| 239 |
+
|
| 240 |
+
name = "mock"
|
| 241 |
+
temperature = TEMPERATURE
|
| 242 |
+
|
| 243 |
+
def __init__(self, model_id: str = MODEL_ID):
|
| 244 |
+
self.model_id = model_id
|
| 245 |
+
# every served model gets its own fabricated numbers; the default keeps the historic ones
|
| 246 |
+
self._salt = "" if model_id == MODEL_ID else f"|{model_id}"
|
| 247 |
+
|
| 248 |
+
def logits(self, state, questions: dict, bias: dict | None = None) -> tuple[dict[str, list[float]], int]:
|
| 249 |
+
evidence = _terms(text(state))
|
| 250 |
+
bias = bias or {}
|
| 251 |
+
out, n_tokens = {}, 0
|
| 252 |
+
for qkey, q in questions.items():
|
| 253 |
+
items = question_options(q)
|
| 254 |
+
labels = LABEL_POOL[: len(items)]
|
| 255 |
+
prompt = SYSTEM + user_message(state, q, labels, items)
|
| 256 |
+
n_tokens += max(1, len(prompt) // 4)
|
| 257 |
+
crit = _terms(text(q.get("instructions")))
|
| 258 |
+
hint = bias.get(qkey) or {}
|
| 259 |
+
row = []
|
| 260 |
+
for i, (key, disp) in enumerate(items):
|
| 261 |
+
opt = _terms(disp)
|
| 262 |
+
overlap = len(opt & evidence) / (len(opt) ** 0.5 + 1.0)
|
| 263 |
+
cue = len(opt & crit) / (len(crit) ** 0.5 + 1.0)
|
| 264 |
+
seed = f"{text(state)}|{text(q.get('instructions'))}|{key}|{i}{self._salt}".encode()
|
| 265 |
+
jitter = int.from_bytes(hashlib.blake2b(seed, digest_size=4).digest(), "big")
|
| 266 |
+
row.append(
|
| 267 |
+
2.6 * overlap
|
| 268 |
+
+ 1.1 * cue
|
| 269 |
+
+ 1.9 * (jitter / 2**32)
|
| 270 |
+
- 0.5 * i / len(items)
|
| 271 |
+
+ float(hint.get(key, 0.0))
|
| 272 |
+
)
|
| 273 |
+
out[qkey] = row
|
| 274 |
+
return out, n_tokens
|
| 275 |
+
|
| 276 |
+
|
| 277 |
+
# --- replay engine ------------------------------------------------------------------
|
| 278 |
+
|
| 279 |
+
|
| 280 |
+
def request_key(state, questions: dict) -> str:
|
| 281 |
+
"""Stable key for a request. Question and option order are part of the request."""
|
| 282 |
+
if isinstance(state, str):
|
| 283 |
+
state = state.replace("\r\n", "\n").strip()
|
| 284 |
+
blob = json.dumps({"state": state, "questions": questions}, ensure_ascii=False, separators=(",", ":"))
|
| 285 |
+
return hashlib.sha256(blob.encode("utf-8")).hexdigest()[:32]
|
| 286 |
+
|
| 287 |
+
|
| 288 |
+
class ReplayEngine:
|
| 289 |
+
"""Label logits the trained model produced for the bundled examples, recorded with
|
| 290 |
+
TorchEngine. Temperature and answer assembly still run live."""
|
| 291 |
+
|
| 292 |
+
name = "replay"
|
| 293 |
+
|
| 294 |
+
def __init__(self, path: str = REPLAY_PATH, temperature: float = TEMPERATURE):
|
| 295 |
+
with open(path, encoding="utf-8") as fh:
|
| 296 |
+
data = json.load(fh)
|
| 297 |
+
self.model_id = data["model"]
|
| 298 |
+
self.hardware = data.get("hardware", "")
|
| 299 |
+
self.temperature = float(temperature)
|
| 300 |
+
self.cache = data["requests"]
|
| 301 |
+
self.last = None
|
| 302 |
+
|
| 303 |
+
def logits(self, state, questions: dict, bias: dict | None = None) -> tuple[dict[str, list[float]], int]:
|
| 304 |
+
del bias
|
| 305 |
+
hit = self.cache.get(request_key(state, questions))
|
| 306 |
+
if hit is None:
|
| 307 |
+
raise ReplayMiss(
|
| 308 |
+
"Edited input. This page only has saved runs for the built-in examples and presets. "
|
| 309 |
+
"Pick one of those, or run blink.py yourself to decide on any text."
|
| 310 |
+
)
|
| 311 |
+
self.last = hit
|
| 312 |
+
return {k: list(v) for k, v in hit["logits"].items()}, int(hit["input_tokens"])
|
| 313 |
+
|
| 314 |
+
|
| 315 |
+
# --- torch engine -------------------------------------------------------------------
|
| 316 |
+
|
| 317 |
+
|
| 318 |
+
def _gpu(duration: int):
|
| 319 |
+
"""spaces.GPU when running on ZeroGPU, a no-op decorator anywhere else."""
|
| 320 |
+
try:
|
| 321 |
+
import spaces
|
| 322 |
+
except Exception:
|
| 323 |
+
return lambda fn: fn
|
| 324 |
+
try:
|
| 325 |
+
return spaces.GPU(duration=duration)
|
| 326 |
+
except Exception:
|
| 327 |
+
return spaces.GPU
|
| 328 |
+
|
| 329 |
+
|
| 330 |
+
def _on_zero_gpu() -> bool:
|
| 331 |
+
try:
|
| 332 |
+
from spaces.config import Config
|
| 333 |
+
|
| 334 |
+
return bool(Config.zero_gpu)
|
| 335 |
+
except Exception:
|
| 336 |
+
return os.environ.get("SPACES_ZERO_GPU", "").lower() in ("1", "true")
|
| 337 |
+
|
| 338 |
+
|
| 339 |
+
# Live engines by id. The GPU function looks its engine up here instead of receiving it as an
|
| 340 |
+
# argument: on ZeroGPU the call runs in a worker process where only module-level state carries the
|
| 341 |
+
# real weights, and an argument would arrive as a copy.
|
| 342 |
+
_LIVE: dict = {}
|
| 343 |
+
|
| 344 |
+
|
| 345 |
+
class TorchEngine:
|
| 346 |
+
"""Prefill-only readout over the offered labels, all questions of a request in one GPU call."""
|
| 347 |
+
|
| 348 |
+
name = "torch"
|
| 349 |
+
|
| 350 |
+
def __init__(self, model_id: str = MODEL_ID, revision=None, temperature: float = TEMPERATURE,
|
| 351 |
+
token_budget: int = TOKEN_BUDGET):
|
| 352 |
+
import torch
|
| 353 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 354 |
+
|
| 355 |
+
self.model_id = model_id
|
| 356 |
+
self.temperature = float(temperature)
|
| 357 |
+
self.token_budget = int(token_budget)
|
| 358 |
+
self.tok = AutoTokenizer.from_pretrained(model_id, revision=revision)
|
| 359 |
+
load = dict(revision=revision, dtype=torch.bfloat16, attn_implementation="sdpa")
|
| 360 |
+
if _on_zero_gpu():
|
| 361 |
+
# ZeroGPU: load on the host, then place on cuda at module level (emulated until
|
| 362 |
+
# a @spaces.GPU call attaches a real device).
|
| 363 |
+
self.model = AutoModelForCausalLM.from_pretrained(model_id, **load).to("cuda")
|
| 364 |
+
else:
|
| 365 |
+
device = "cuda" if torch.cuda.is_available() else "cpu"
|
| 366 |
+
self.model = AutoModelForCausalLM.from_pretrained(model_id, device_map=device, **load)
|
| 367 |
+
self.model.eval()
|
| 368 |
+
self.pad_id = self.tok.pad_token_id if self.tok.pad_token_id is not None else 0
|
| 369 |
+
self.labels, self.label_ids = self._verify_labels()
|
| 370 |
+
self.key = id(self)
|
| 371 |
+
_LIVE[self.key] = self
|
| 372 |
+
|
| 373 |
+
def _wrap(self, user: str) -> str:
|
| 374 |
+
return self.tok.apply_chat_template(
|
| 375 |
+
[{"role": "system", "content": SYSTEM}, {"role": "user", "content": user}],
|
| 376 |
+
tokenize=False,
|
| 377 |
+
add_generation_prompt=True,
|
| 378 |
+
enable_thinking=False,
|
| 379 |
+
)
|
| 380 |
+
|
| 381 |
+
def _verify_labels(self) -> tuple[list[str], list[int]]:
|
| 382 |
+
"""Keep labels that are one token at the assistant boundary and decode back to themselves."""
|
| 383 |
+
probe = self._wrap("x")
|
| 384 |
+
base = self.tok(probe, add_special_tokens=False)["input_ids"]
|
| 385 |
+
labels, ids = [], []
|
| 386 |
+
for lab in LABEL_POOL:
|
| 387 |
+
t = self.tok(probe + lab, add_special_tokens=False)["input_ids"]
|
| 388 |
+
if (
|
| 389 |
+
len(t) == len(base) + 1
|
| 390 |
+
and t[: len(base)] == base
|
| 391 |
+
and self.tok.decode(t[-1:]) == lab
|
| 392 |
+
):
|
| 393 |
+
labels.append(lab)
|
| 394 |
+
ids.append(t[-1])
|
| 395 |
+
if len(labels) >= MAX_OPTIONS:
|
| 396 |
+
break
|
| 397 |
+
if len(set(ids)) != len(ids):
|
| 398 |
+
raise BlinkError("label token ids collide")
|
| 399 |
+
if len(labels) < 26 or labels[:26] != list(string.ascii_uppercase):
|
| 400 |
+
raise BlinkError(f"only {len(labels)} verified single-token labels")
|
| 401 |
+
return labels, ids
|
| 402 |
+
|
| 403 |
+
def render(self, state, questions: dict):
|
| 404 |
+
work = []
|
| 405 |
+
for qkey, q in questions.items():
|
| 406 |
+
items = question_options(q)
|
| 407 |
+
if len(items) > len(self.labels):
|
| 408 |
+
raise BlinkError(f"{len(items)} options per choice; {len(self.labels)} labels verified")
|
| 409 |
+
labels = self.labels[: len(items)]
|
| 410 |
+
prompt = self._wrap(user_message(state, q, labels, items))
|
| 411 |
+
ids = self.tok(prompt, add_special_tokens=False)["input_ids"]
|
| 412 |
+
if len(ids) > MAX_INPUT_TOKENS:
|
| 413 |
+
raise BlinkError(
|
| 414 |
+
f"question {qkey!r} renders to {len(ids)} tokens, over the maximum context length "
|
| 415 |
+
f"of {MAX_INPUT_TOKENS}"
|
| 416 |
+
)
|
| 417 |
+
work.append(
|
| 418 |
+
{
|
| 419 |
+
"qkey": qkey,
|
| 420 |
+
"keys": [k for k, _ in items],
|
| 421 |
+
"ids": ids,
|
| 422 |
+
"cand": self.label_ids[: len(items)],
|
| 423 |
+
}
|
| 424 |
+
)
|
| 425 |
+
return work
|
| 426 |
+
|
| 427 |
+
def logits(self, state, questions: dict, bias: dict | None = None) -> tuple[dict[str, list[float]], int]:
|
| 428 |
+
del bias # demo-only shaping; the trained model reads the evidence instead
|
| 429 |
+
work = self.render(state, questions)
|
| 430 |
+
rows, self.last_model_ms = _forward(self.key, [w["ids"] for w in work], [w["cand"] for w in work])
|
| 431 |
+
return (
|
| 432 |
+
{w["qkey"]: r for w, r in zip(work, rows)},
|
| 433 |
+
sum(len(w["ids"]) for w in work),
|
| 434 |
+
)
|
| 435 |
+
|
| 436 |
+
|
| 437 |
+
def _batches(lengths: list[int], budget: int):
|
| 438 |
+
"""Shortest first; a batch closes when (longest x count) would pass the padded-token budget.
|
| 439 |
+
A sequence longer than the budget runs alone."""
|
| 440 |
+
order = sorted(range(len(lengths)), key=lambda i: lengths[i])
|
| 441 |
+
batch, longest = [], 0
|
| 442 |
+
for i in order:
|
| 443 |
+
grown = max(longest, lengths[i])
|
| 444 |
+
if batch and grown * (len(batch) + 1) > budget:
|
| 445 |
+
yield batch
|
| 446 |
+
batch, grown = [], lengths[i]
|
| 447 |
+
batch.append(i)
|
| 448 |
+
longest = grown
|
| 449 |
+
if batch:
|
| 450 |
+
yield batch
|
| 451 |
+
|
| 452 |
+
|
| 453 |
+
@_gpu(GPU_DURATION)
|
| 454 |
+
def _forward(key: int, seqs: list[list[int]], cands: list[list[int]]):
|
| 455 |
+
"""Right-padded maskless prefill; every layer is causal, so padding cannot reach the
|
| 456 |
+
last real position. Label rows of lm_head are applied in FP32."""
|
| 457 |
+
import torch
|
| 458 |
+
|
| 459 |
+
engine = _LIVE[key]
|
| 460 |
+
model = engine.model
|
| 461 |
+
device = next(model.parameters()).device
|
| 462 |
+
head = model.lm_head.weight
|
| 463 |
+
out: list = [None] * len(seqs)
|
| 464 |
+
t0 = time.perf_counter()
|
| 465 |
+
with torch.no_grad():
|
| 466 |
+
for b in _batches([len(s) for s in seqs], getattr(engine, "token_budget", TOKEN_BUDGET)):
|
| 467 |
+
L = max(len(seqs[i]) for i in b)
|
| 468 |
+
ids = torch.full((len(b), L), engine.pad_id, dtype=torch.long)
|
| 469 |
+
for r, i in enumerate(b):
|
| 470 |
+
ids[r, : len(seqs[i])] = torch.tensor(seqs[i], dtype=torch.long)
|
| 471 |
+
ids = ids.to(device)
|
| 472 |
+
h = model.model(input_ids=ids, use_cache=False).last_hidden_state
|
| 473 |
+
last = torch.tensor([len(seqs[i]) - 1 for i in b], device=device)
|
| 474 |
+
h = h[torch.arange(len(b), device=device), last].float()
|
| 475 |
+
for r, i in enumerate(b):
|
| 476 |
+
w = head[torch.tensor(cands[i], device=device)].float()
|
| 477 |
+
out[i] = (w @ h[r]).tolist()
|
| 478 |
+
return out, round((time.perf_counter() - t0) * 1000, 1)
|
| 479 |
+
|
| 480 |
+
|
| 481 |
+
# --- hybrid engine ------------------------------------------------------------------
|
| 482 |
+
|
| 483 |
+
|
| 484 |
+
class HybridEngine:
|
| 485 |
+
"""Recorded outputs for the bundled examples (instant, no GPU), the live model for
|
| 486 |
+
everything else. The torch engine is built eagerly: ZeroGPU wants weights placed at startup."""
|
| 487 |
+
|
| 488 |
+
name = "hybrid"
|
| 489 |
+
|
| 490 |
+
def __init__(self, replay: "ReplayEngine | None", live: "TorchEngine"):
|
| 491 |
+
self.replay, self.live = replay, live
|
| 492 |
+
self.model_id = live.model_id
|
| 493 |
+
self.temperature = live.temperature
|
| 494 |
+
self.hardware = replay.hardware if replay is not None else ""
|
| 495 |
+
self.last_source = None
|
| 496 |
+
self.last = None
|
| 497 |
+
|
| 498 |
+
def logits(self, state, questions: dict, bias: dict | None = None,
|
| 499 |
+
prefer: str = "live") -> tuple[dict[str, list[float]], int]:
|
| 500 |
+
if (
|
| 501 |
+
prefer == "saved"
|
| 502 |
+
and self.replay is not None
|
| 503 |
+
and self.replay.cache.get(request_key(state, questions)) is not None
|
| 504 |
+
):
|
| 505 |
+
self.last_source = "replay"
|
| 506 |
+
out = self.replay.logits(state, questions)
|
| 507 |
+
self.last = self.replay.last
|
| 508 |
+
return out
|
| 509 |
+
self.last_source = "torch"
|
| 510 |
+
return self.live.logits(state, questions)
|
| 511 |
+
|
| 512 |
+
|
| 513 |
+
# --- public entry point -------------------------------------------------------------
|
| 514 |
+
|
| 515 |
+
_ENGINE = None # the default model's engine (tests inject one here)
|
| 516 |
+
_ENGINES: dict = {} # every other served model's engine, by id
|
| 517 |
+
|
| 518 |
+
|
| 519 |
+
def _cuda() -> bool:
|
| 520 |
+
try:
|
| 521 |
+
import torch
|
| 522 |
+
except Exception:
|
| 523 |
+
return False
|
| 524 |
+
return bool(torch.cuda.is_available())
|
| 525 |
+
|
| 526 |
+
|
| 527 |
+
def engine_kind() -> str:
|
| 528 |
+
if ENGINE in ("mock", "replay", "torch", "hybrid"):
|
| 529 |
+
return ENGINE
|
| 530 |
+
recorded = os.path.exists(REPLAY_PATH)
|
| 531 |
+
if _cuda():
|
| 532 |
+
return "hybrid" if recorded else "torch"
|
| 533 |
+
return "replay" if recorded else "torch"
|
| 534 |
+
|
| 535 |
+
|
| 536 |
+
def models() -> list[str]:
|
| 537 |
+
"""Model ids this deployment serves, the default first."""
|
| 538 |
+
return [m for m, _ in MODEL_SPECS]
|
| 539 |
+
|
| 540 |
+
|
| 541 |
+
def _matching_replay(model_id: str):
|
| 542 |
+
"""The model's recording, only if that same model made it: a saved run must be what the live
|
| 543 |
+
model would answer."""
|
| 544 |
+
path = replay_path(model_id)
|
| 545 |
+
if not os.path.exists(path):
|
| 546 |
+
return None
|
| 547 |
+
rec = ReplayEngine(path, TEMPERATURE)
|
| 548 |
+
return rec if rec.model_id == model_id else None
|
| 549 |
+
|
| 550 |
+
|
| 551 |
+
def _build(model_id: str):
|
| 552 |
+
revision = dict(MODEL_SPECS).get(model_id)
|
| 553 |
+
kind = engine_kind()
|
| 554 |
+
if kind == "mock":
|
| 555 |
+
return MockEngine(model_id)
|
| 556 |
+
if kind == "replay":
|
| 557 |
+
return ReplayEngine(replay_path(model_id), TEMPERATURE)
|
| 558 |
+
if kind == "hybrid":
|
| 559 |
+
return HybridEngine(_matching_replay(model_id), TorchEngine(model_id, revision, TEMPERATURE))
|
| 560 |
+
return TorchEngine(model_id, revision, TEMPERATURE)
|
| 561 |
+
|
| 562 |
+
|
| 563 |
+
def engine(model: str | None = None):
|
| 564 |
+
global _ENGINE
|
| 565 |
+
mid = model or MODEL_ID
|
| 566 |
+
if mid not in dict(MODEL_SPECS):
|
| 567 |
+
raise BlinkError(f"unknown model {mid!r}; this deployment serves {', '.join(models())}")
|
| 568 |
+
if mid == MODEL_ID:
|
| 569 |
+
if _ENGINE is None:
|
| 570 |
+
_ENGINE = _build(mid)
|
| 571 |
+
return _ENGINE
|
| 572 |
+
if mid not in _ENGINES:
|
| 573 |
+
_ENGINES[mid] = _build(mid)
|
| 574 |
+
return _ENGINES[mid]
|
| 575 |
+
|
| 576 |
+
|
| 577 |
+
def warm() -> list:
|
| 578 |
+
"""Build every served model's engine now: ZeroGPU wants weights placed at startup."""
|
| 579 |
+
return [engine(m) for m in models()]
|
| 580 |
+
|
| 581 |
+
|
| 582 |
+
def decide(state, questions: dict, temperature: float | None = None, bias: dict | None = None,
|
| 583 |
+
prefer: str = "live", model: str | None = None) -> dict:
|
| 584 |
+
"""{state, questions} -> {answers, meta}. One forward pass, zero generated tokens.
|
| 585 |
+
|
| 586 |
+
`bias` shapes the mock engine so the bundled examples read realistically without a
|
| 587 |
+
GPU. It is discarded by the trained model. `prefer="saved"` lets the hybrid engine answer
|
| 588 |
+
a bundled example from its recording (used for the first page render); every other call
|
| 589 |
+
runs the model. `model` picks one of models() (default: the first).
|
| 590 |
+
"""
|
| 591 |
+
validate(questions)
|
| 592 |
+
eng = engine(model)
|
| 593 |
+
T = float(temperature if temperature is not None else eng.temperature)
|
| 594 |
+
if T <= 0:
|
| 595 |
+
raise BlinkError("temperature must be positive")
|
| 596 |
+
t0 = time.perf_counter()
|
| 597 |
+
if isinstance(eng, HybridEngine):
|
| 598 |
+
raw, n_tokens = eng.logits(state, questions, bias, prefer=prefer)
|
| 599 |
+
else:
|
| 600 |
+
raw, n_tokens = eng.logits(state, questions, bias)
|
| 601 |
+
answers = {}
|
| 602 |
+
for qkey, q in questions.items():
|
| 603 |
+
keys = [k for k, _ in question_options(q)]
|
| 604 |
+
answers[qkey] = answer_for(q, keys, softmax(raw[qkey], T))
|
| 605 |
+
latency_ms = round((time.perf_counter() - t0) * 1000, 1)
|
| 606 |
+
source = getattr(eng, "last_source", None) or eng.name
|
| 607 |
+
meta = {
|
| 608 |
+
"model": getattr(eng, "model_id", MODEL_ID),
|
| 609 |
+
"engine": source,
|
| 610 |
+
"temperature": T,
|
| 611 |
+
"input_tokens": n_tokens,
|
| 612 |
+
"generated_tokens": 0,
|
| 613 |
+
"latency_ms": latency_ms,
|
| 614 |
+
}
|
| 615 |
+
if source == "replay":
|
| 616 |
+
meta["latency_ms"] = float(eng.last["latency_ms"])
|
| 617 |
+
if eng.hardware:
|
| 618 |
+
meta["recorded_on"] = eng.hardware
|
| 619 |
+
elif source == "torch":
|
| 620 |
+
model_ms = getattr(getattr(eng, "live", eng), "last_model_ms", None)
|
| 621 |
+
if model_ms is not None:
|
| 622 |
+
meta["model_ms"] = model_ms # forward passes only; excludes any wait for a device
|
| 623 |
+
return {"answers": answers, "meta": meta}
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- macro render_value(value) -%}
|
| 2 |
+
{%- if value is string -%}
|
| 3 |
+
{{- value -}}
|
| 4 |
+
{%- else -%}
|
| 5 |
+
{{- value | tojson(ensure_ascii=False) -}}
|
| 6 |
+
{%- endif -%}
|
| 7 |
+
{%- endmacro -%}
|
| 8 |
+
|
| 9 |
+
{%- macro render_content(message_content) -%}
|
| 10 |
+
{%- if message_content is string -%}
|
| 11 |
+
{{- message_content -}}
|
| 12 |
+
{%- elif message_content is iterable -%}
|
| 13 |
+
{%- for part in message_content -%}
|
| 14 |
+
{%- if part is not mapping -%}
|
| 15 |
+
{{- part -}}
|
| 16 |
+
{%- elif part['type'] == 'image' or 'image' in part or 'image_url' in part -%}
|
| 17 |
+
{{- '<|vision_start|><|image_pad|><|vision_end|>' -}}
|
| 18 |
+
{%- elif part['type'] == 'audio' or part['type'] == 'input_audio' or 'audio' in part or 'audio_url' in part or 'input_audio' in part -%}
|
| 19 |
+
{{- '<|mimo_audio_start|><|audio_pad|><|mimo_audio_end|>' -}}
|
| 20 |
+
{%- elif part['type'] == 'video' or 'video' in part or 'video_url' in part -%}
|
| 21 |
+
{{- '<|vision_start|><|video_pad|><|vision_end|>' -}}
|
| 22 |
+
{%- elif 'text' in part -%}
|
| 23 |
+
{{- part['text'] -}}
|
| 24 |
+
{%- endif -%}
|
| 25 |
+
{%- endfor -%}
|
| 26 |
+
{%- endif -%}
|
| 27 |
+
{%- endmacro -%}
|
| 28 |
+
|
| 29 |
+
{%- macro render_tools(tools) -%}
|
| 30 |
+
{{- 'You are provided with the following tools:\n\n<tools>' -}}
|
| 31 |
+
{%- for tool in tools -%}
|
| 32 |
+
{{- '\n' ~ (tool | tojson(ensure_ascii=False)) -}}
|
| 33 |
+
{%- endfor -%}
|
| 34 |
+
{{- '\n</tools>' -}}
|
| 35 |
+
{%- endmacro -%}
|
| 36 |
+
|
| 37 |
+
{%- macro render_tool_calls(tool_calls) -%}
|
| 38 |
+
{%- for tool_call in tool_calls -%}
|
| 39 |
+
{%- if tool_call.function is defined -%}
|
| 40 |
+
{%- set tool_call = tool_call.function -%}
|
| 41 |
+
{%- elif tool_call.custom is defined -%}
|
| 42 |
+
{%- set tool_call = tool_call.custom -%}
|
| 43 |
+
{%- endif -%}
|
| 44 |
+
{{- '<tool_call><function=' ~ tool_call.name ~ '>' -}}
|
| 45 |
+
{%- if tool_call.input is defined and tool_call.input is string -%}
|
| 46 |
+
{{- tool_call.input -}}
|
| 47 |
+
{%- elif tool_call.arguments -%}
|
| 48 |
+
{%- if tool_call.arguments is string -%}
|
| 49 |
+
{{- tool_call.arguments -}}
|
| 50 |
+
{%- else -%}
|
| 51 |
+
{%- for args_name, args_value in tool_call.arguments | items -%}
|
| 52 |
+
{{- '<parameter=' ~ args_name ~ '>' ~ render_value(args_value) ~ '</parameter>' -}}
|
| 53 |
+
{%- endfor -%}
|
| 54 |
+
{%- endif -%}
|
| 55 |
+
{%- endif -%}
|
| 56 |
+
{{- '</function></tool_call>' -}}
|
| 57 |
+
{%- endfor -%}
|
| 58 |
+
{%- endmacro -%}
|
| 59 |
+
|
| 60 |
+
{%- macro render_assistant_message(message) -%}
|
| 61 |
+
{%- generation -%}
|
| 62 |
+
{%- set content = render_content(message.content) -%}
|
| 63 |
+
{%- set reasoning = message.reasoning_content if message.reasoning_content is string else '' -%}
|
| 64 |
+
{{- '<|im_start|>assistant\n<think>' ~ reasoning ~ '</think>' ~ content -}}
|
| 65 |
+
{%- if message.tool_calls is defined and message.tool_calls is iterable and message.tool_calls | length > 0 -%}
|
| 66 |
+
{{- render_tool_calls(message.tool_calls) -}}
|
| 67 |
+
{%- endif -%}
|
| 68 |
+
{{- '<|im_end|>' -}}
|
| 69 |
+
{%- endgeneration -%}
|
| 70 |
+
{%- endmacro -%}
|
| 71 |
+
|
| 72 |
+
{%- if tools is defined and tools is iterable and tools | length > 0 -%}
|
| 73 |
+
{{- '<|im_start|>system\n' ~ render_tools(tools) ~ '<|im_end|>' -}}
|
| 74 |
+
{%- endif -%}
|
| 75 |
+
|
| 76 |
+
{%- for message in messages -%}
|
| 77 |
+
{%- if message.role == 'assistant' -%}
|
| 78 |
+
{{- render_assistant_message(message) -}}
|
| 79 |
+
{%- else -%}
|
| 80 |
+
{%- set body = render_content(message.content) -%}
|
| 81 |
+
{{- '<|im_start|>' ~ message.role ~ '\n' ~ body -}}
|
| 82 |
+
{%- if message.tools is defined and message.tools is iterable and message.tools | length > 0 -%}
|
| 83 |
+
{%- if body -%}
|
| 84 |
+
{{- '\n\n' -}}
|
| 85 |
+
{%- endif -%}
|
| 86 |
+
{{- render_tools(message.tools) -}}
|
| 87 |
+
{%- endif -%}
|
| 88 |
+
{{- '<|im_end|>' -}}
|
| 89 |
+
{%- endif -%}
|
| 90 |
+
{%- endfor -%}
|
| 91 |
+
|
| 92 |
+
{%- if add_generation_prompt -%}
|
| 93 |
+
{{- '<|im_start|>assistant\n' -}}
|
| 94 |
+
{%- if enable_thinking is false -%}
|
| 95 |
+
{{- '<think></think>' -}}
|
| 96 |
+
{%- endif -%}
|
| 97 |
+
{%- endif -%}
|
config.json
ADDED
|
@@ -0,0 +1,107 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"Qwen3_5ForConditionalGeneration"
|
| 4 |
+
],
|
| 5 |
+
"image_token_id": 248056,
|
| 6 |
+
"model_type": "qwen3_5",
|
| 7 |
+
"text_config": {
|
| 8 |
+
"attention_bias": false,
|
| 9 |
+
"attention_dropout": 0.0,
|
| 10 |
+
"attn_output_gate": true,
|
| 11 |
+
"bos_token_id": null,
|
| 12 |
+
"dtype": "bfloat16",
|
| 13 |
+
"eos_token_id": 248044,
|
| 14 |
+
"full_attention_interval": 4,
|
| 15 |
+
"head_dim": 256,
|
| 16 |
+
"hidden_act": "silu",
|
| 17 |
+
"hidden_size": 4096,
|
| 18 |
+
"initializer_range": 0.02,
|
| 19 |
+
"intermediate_size": 12288,
|
| 20 |
+
"layer_types": [
|
| 21 |
+
"linear_attention",
|
| 22 |
+
"linear_attention",
|
| 23 |
+
"linear_attention",
|
| 24 |
+
"full_attention",
|
| 25 |
+
"linear_attention",
|
| 26 |
+
"linear_attention",
|
| 27 |
+
"linear_attention",
|
| 28 |
+
"full_attention",
|
| 29 |
+
"linear_attention",
|
| 30 |
+
"linear_attention",
|
| 31 |
+
"linear_attention",
|
| 32 |
+
"full_attention",
|
| 33 |
+
"linear_attention",
|
| 34 |
+
"linear_attention",
|
| 35 |
+
"linear_attention",
|
| 36 |
+
"full_attention",
|
| 37 |
+
"linear_attention",
|
| 38 |
+
"linear_attention",
|
| 39 |
+
"linear_attention",
|
| 40 |
+
"full_attention",
|
| 41 |
+
"linear_attention",
|
| 42 |
+
"linear_attention",
|
| 43 |
+
"linear_attention",
|
| 44 |
+
"full_attention",
|
| 45 |
+
"linear_attention",
|
| 46 |
+
"linear_attention",
|
| 47 |
+
"linear_attention",
|
| 48 |
+
"full_attention",
|
| 49 |
+
"linear_attention",
|
| 50 |
+
"linear_attention",
|
| 51 |
+
"linear_attention",
|
| 52 |
+
"full_attention"
|
| 53 |
+
],
|
| 54 |
+
"linear_conv_kernel_dim": 4,
|
| 55 |
+
"linear_key_head_dim": 128,
|
| 56 |
+
"linear_num_key_heads": 16,
|
| 57 |
+
"linear_num_value_heads": 32,
|
| 58 |
+
"linear_value_head_dim": 128,
|
| 59 |
+
"mamba_ssm_dtype": "float32",
|
| 60 |
+
"max_position_embeddings": 262144,
|
| 61 |
+
"mlp_only_layers": [],
|
| 62 |
+
"model_type": "qwen3_5_text",
|
| 63 |
+
"mtp_num_hidden_layers": 1,
|
| 64 |
+
"mtp_use_dedicated_embeddings": false,
|
| 65 |
+
"num_attention_heads": 16,
|
| 66 |
+
"num_hidden_layers": 32,
|
| 67 |
+
"num_key_value_heads": 4,
|
| 68 |
+
"pad_token_id": null,
|
| 69 |
+
"partial_rotary_factor": 0.25,
|
| 70 |
+
"rms_norm_eps": 1e-06,
|
| 71 |
+
"rope_parameters": {
|
| 72 |
+
"mrope_interleaved": true,
|
| 73 |
+
"mrope_section": [
|
| 74 |
+
11,
|
| 75 |
+
11,
|
| 76 |
+
10
|
| 77 |
+
],
|
| 78 |
+
"partial_rotary_factor": 0.25,
|
| 79 |
+
"rope_theta": 10000000,
|
| 80 |
+
"rope_type": "default"
|
| 81 |
+
},
|
| 82 |
+
"tie_word_embeddings": false,
|
| 83 |
+
"use_cache": true,
|
| 84 |
+
"vocab_size": 248320
|
| 85 |
+
},
|
| 86 |
+
"tie_word_embeddings": false,
|
| 87 |
+
"transformers_version": "5.12.1",
|
| 88 |
+
"video_token_id": 248057,
|
| 89 |
+
"vision_config": {
|
| 90 |
+
"deepstack_visual_indexes": [],
|
| 91 |
+
"depth": 27,
|
| 92 |
+
"hidden_act": "gelu_pytorch_tanh",
|
| 93 |
+
"hidden_size": 1152,
|
| 94 |
+
"in_channels": 3,
|
| 95 |
+
"initializer_range": 0.02,
|
| 96 |
+
"intermediate_size": 4304,
|
| 97 |
+
"model_type": "qwen3_5_vision",
|
| 98 |
+
"num_heads": 16,
|
| 99 |
+
"num_position_embeddings": 2304,
|
| 100 |
+
"out_hidden_size": 4096,
|
| 101 |
+
"patch_size": 16,
|
| 102 |
+
"spatial_merge_size": 2,
|
| 103 |
+
"temporal_patch_size": 2
|
| 104 |
+
},
|
| 105 |
+
"vision_end_token_id": 248054,
|
| 106 |
+
"vision_start_token_id": 248053
|
| 107 |
+
}
|
eval/decision-index-0.1-full.json
ADDED
|
@@ -0,0 +1,1080 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"engine": "blink-mimo-9b",
|
| 3 |
+
"generated_utc": "2026-09-24T14:47:48+00:00",
|
| 4 |
+
"suite": {
|
| 5 |
+
"requests": 132422,
|
| 6 |
+
"scoreable": 131980,
|
| 7 |
+
"excluded": 442,
|
| 8 |
+
"benchmarks": 37,
|
| 9 |
+
"rows_sha256": "750d353a3a83af615c67cfe9752e005bf09e6c28d9c4ba28d3a9f57ba8536cfd"
|
| 10 |
+
},
|
| 11 |
+
"completed": 132422,
|
| 12 |
+
"complete": true,
|
| 13 |
+
"counts": {
|
| 14 |
+
"ok": 132422
|
| 15 |
+
},
|
| 16 |
+
"latency_ms": {
|
| 17 |
+
"median": 49.9,
|
| 18 |
+
"p95": 1129.5,
|
| 19 |
+
"mean": 251.2
|
| 20 |
+
},
|
| 21 |
+
"decision_index": 56.53,
|
| 22 |
+
"scores": {
|
| 23 |
+
"balanced_raw": 56.53,
|
| 24 |
+
"balanced_skill": 42.17,
|
| 25 |
+
"breadth_skill": 40.37
|
| 26 |
+
},
|
| 27 |
+
"areas": [
|
| 28 |
+
{
|
| 29 |
+
"id": "knowledge",
|
| 30 |
+
"label": "Knowledge & Reasoning",
|
| 31 |
+
"raw": 0.5511,
|
| 32 |
+
"skill": 0.3834,
|
| 33 |
+
"coverage": 1.0,
|
| 34 |
+
"pending": 0.0,
|
| 35 |
+
"n": 6,
|
| 36 |
+
"benchmarks": [
|
| 37 |
+
24,
|
| 38 |
+
25,
|
| 39 |
+
30,
|
| 40 |
+
43,
|
| 41 |
+
44,
|
| 42 |
+
31
|
| 43 |
+
]
|
| 44 |
+
},
|
| 45 |
+
{
|
| 46 |
+
"id": "language",
|
| 47 |
+
"label": "Language Understanding",
|
| 48 |
+
"raw": 0.7011,
|
| 49 |
+
"skill": 0.5901,
|
| 50 |
+
"coverage": 1.0,
|
| 51 |
+
"pending": 0.0,
|
| 52 |
+
"n": 3,
|
| 53 |
+
"benchmarks": [
|
| 54 |
+
11,
|
| 55 |
+
40,
|
| 56 |
+
41
|
| 57 |
+
]
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"id": "retrieval",
|
| 61 |
+
"label": "Retrieval & Classification",
|
| 62 |
+
"raw": 0.3484,
|
| 63 |
+
"skill": 0.2683,
|
| 64 |
+
"coverage": 1.0,
|
| 65 |
+
"pending": 0.0,
|
| 66 |
+
"n": 2,
|
| 67 |
+
"benchmarks": [
|
| 68 |
+
36,
|
| 69 |
+
37
|
| 70 |
+
]
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"id": "tools",
|
| 74 |
+
"label": "Tools & Automation",
|
| 75 |
+
"raw": 0.7045,
|
| 76 |
+
"skill": 0.5811,
|
| 77 |
+
"coverage": 1.0,
|
| 78 |
+
"pending": 0.0,
|
| 79 |
+
"n": 3,
|
| 80 |
+
"benchmarks": [
|
| 81 |
+
1,
|
| 82 |
+
2,
|
| 83 |
+
6
|
| 84 |
+
]
|
| 85 |
+
},
|
| 86 |
+
{
|
| 87 |
+
"id": "arts",
|
| 88 |
+
"label": "Arts & Human Judgment",
|
| 89 |
+
"raw": 0.5214,
|
| 90 |
+
"skill": 0.2859,
|
| 91 |
+
"coverage": 1.0,
|
| 92 |
+
"pending": 0.0,
|
| 93 |
+
"n": 5,
|
| 94 |
+
"benchmarks": [
|
| 95 |
+
20,
|
| 96 |
+
21,
|
| 97 |
+
22,
|
| 98 |
+
23,
|
| 99 |
+
50
|
| 100 |
+
]
|
| 101 |
+
}
|
| 102 |
+
],
|
| 103 |
+
"index_benchmarks": {
|
| 104 |
+
"1": {
|
| 105 |
+
"raw": 0.8926,
|
| 106 |
+
"skill": 0.855,
|
| 107 |
+
"coverage": 1.0,
|
| 108 |
+
"pending": 0.0,
|
| 109 |
+
"random": 0.2592,
|
| 110 |
+
"tracks": [],
|
| 111 |
+
"in_index": true
|
| 112 |
+
},
|
| 113 |
+
"2": {
|
| 114 |
+
"raw": 0.4222,
|
| 115 |
+
"skill": 0.3627,
|
| 116 |
+
"coverage": 1.0,
|
| 117 |
+
"pending": 0.0,
|
| 118 |
+
"random": 0.0933,
|
| 119 |
+
"tracks": [],
|
| 120 |
+
"in_index": true
|
| 121 |
+
},
|
| 122 |
+
"6": {
|
| 123 |
+
"raw": 0.7989,
|
| 124 |
+
"skill": 0.5256,
|
| 125 |
+
"coverage": 1.0,
|
| 126 |
+
"pending": 0.0,
|
| 127 |
+
"random": 0.5245,
|
| 128 |
+
"tracks": [
|
| 129 |
+
{
|
| 130 |
+
"track": "RouterBench-0shot",
|
| 131 |
+
"score": 0.7864,
|
| 132 |
+
"headline": false
|
| 133 |
+
},
|
| 134 |
+
{
|
| 135 |
+
"track": "RouterBench-5shot",
|
| 136 |
+
"score": 0.8114,
|
| 137 |
+
"headline": false
|
| 138 |
+
}
|
| 139 |
+
],
|
| 140 |
+
"in_index": true
|
| 141 |
+
},
|
| 142 |
+
"7": {
|
| 143 |
+
"raw": 0.0,
|
| 144 |
+
"skill": 0.0,
|
| 145 |
+
"coverage": 0.0,
|
| 146 |
+
"pending": 1.0,
|
| 147 |
+
"random": null,
|
| 148 |
+
"tracks": [],
|
| 149 |
+
"in_index": false
|
| 150 |
+
},
|
| 151 |
+
"11": {
|
| 152 |
+
"raw": 0.8171,
|
| 153 |
+
"skill": 0.7355,
|
| 154 |
+
"coverage": 1.0,
|
| 155 |
+
"pending": 0.0,
|
| 156 |
+
"random": 0.3085,
|
| 157 |
+
"tracks": [],
|
| 158 |
+
"in_index": true
|
| 159 |
+
},
|
| 160 |
+
"13": {
|
| 161 |
+
"raw": 0.0,
|
| 162 |
+
"skill": 0.0,
|
| 163 |
+
"coverage": 0.0,
|
| 164 |
+
"pending": 1.0,
|
| 165 |
+
"random": null,
|
| 166 |
+
"tracks": [],
|
| 167 |
+
"in_index": false
|
| 168 |
+
},
|
| 169 |
+
"15": {
|
| 170 |
+
"raw": 0.0,
|
| 171 |
+
"skill": 0.0,
|
| 172 |
+
"coverage": 0.0,
|
| 173 |
+
"pending": 1.0,
|
| 174 |
+
"random": null,
|
| 175 |
+
"tracks": [],
|
| 176 |
+
"in_index": false
|
| 177 |
+
},
|
| 178 |
+
"16": {
|
| 179 |
+
"raw": 0.0,
|
| 180 |
+
"skill": 0.0,
|
| 181 |
+
"coverage": 0.0,
|
| 182 |
+
"pending": 1.0,
|
| 183 |
+
"random": null,
|
| 184 |
+
"tracks": [],
|
| 185 |
+
"in_index": false
|
| 186 |
+
},
|
| 187 |
+
"17": {
|
| 188 |
+
"raw": 0.0,
|
| 189 |
+
"skill": 0.0,
|
| 190 |
+
"coverage": 0.0,
|
| 191 |
+
"pending": 1.0,
|
| 192 |
+
"random": null,
|
| 193 |
+
"tracks": [],
|
| 194 |
+
"in_index": false
|
| 195 |
+
},
|
| 196 |
+
"19": {
|
| 197 |
+
"raw": 0.0,
|
| 198 |
+
"skill": 0.0,
|
| 199 |
+
"coverage": 0.0,
|
| 200 |
+
"pending": 1.0,
|
| 201 |
+
"random": null,
|
| 202 |
+
"tracks": [],
|
| 203 |
+
"in_index": false
|
| 204 |
+
},
|
| 205 |
+
"20": {
|
| 206 |
+
"raw": 0.8408,
|
| 207 |
+
"skill": 0.6817,
|
| 208 |
+
"coverage": 1.0,
|
| 209 |
+
"pending": 0.0,
|
| 210 |
+
"random": 0.5,
|
| 211 |
+
"tracks": [],
|
| 212 |
+
"in_index": true
|
| 213 |
+
},
|
| 214 |
+
"21": {
|
| 215 |
+
"raw": 0.6381,
|
| 216 |
+
"skill": 0.2763,
|
| 217 |
+
"coverage": 1.0,
|
| 218 |
+
"pending": 0.0,
|
| 219 |
+
"random": 0.5,
|
| 220 |
+
"tracks": [],
|
| 221 |
+
"in_index": true
|
| 222 |
+
},
|
| 223 |
+
"22": {
|
| 224 |
+
"raw": 0.0763,
|
| 225 |
+
"skill": 0.0691,
|
| 226 |
+
"coverage": 1.0,
|
| 227 |
+
"pending": 0.0,
|
| 228 |
+
"random": 0.0078,
|
| 229 |
+
"tracks": [],
|
| 230 |
+
"in_index": true
|
| 231 |
+
},
|
| 232 |
+
"23": {
|
| 233 |
+
"raw": 0.5965,
|
| 234 |
+
"skill": 0.193,
|
| 235 |
+
"coverage": 1.0,
|
| 236 |
+
"pending": 0.0,
|
| 237 |
+
"random": 0.5,
|
| 238 |
+
"tracks": [],
|
| 239 |
+
"in_index": true
|
| 240 |
+
},
|
| 241 |
+
"24": {
|
| 242 |
+
"raw": 0.8019,
|
| 243 |
+
"skill": 0.7359,
|
| 244 |
+
"coverage": 1.0,
|
| 245 |
+
"pending": 0.0,
|
| 246 |
+
"random": 0.25,
|
| 247 |
+
"tracks": [],
|
| 248 |
+
"in_index": true
|
| 249 |
+
},
|
| 250 |
+
"25": {
|
| 251 |
+
"raw": 0.4082,
|
| 252 |
+
"skill": 0.2109,
|
| 253 |
+
"coverage": 1.0,
|
| 254 |
+
"pending": 0.0,
|
| 255 |
+
"random": 0.25,
|
| 256 |
+
"tracks": [],
|
| 257 |
+
"in_index": true
|
| 258 |
+
},
|
| 259 |
+
"30": {
|
| 260 |
+
"raw": 0.6585,
|
| 261 |
+
"skill": 0.5882,
|
| 262 |
+
"coverage": 1.0,
|
| 263 |
+
"pending": 0.0,
|
| 264 |
+
"random": 0.25,
|
| 265 |
+
"tracks": [
|
| 266 |
+
{
|
| 267 |
+
"track": "GSM8K-4choice",
|
| 268 |
+
"score": 0.7096,
|
| 269 |
+
"headline": false
|
| 270 |
+
},
|
| 271 |
+
{
|
| 272 |
+
"track": "GSM8K-10choice",
|
| 273 |
+
"score": 0.6073,
|
| 274 |
+
"headline": false
|
| 275 |
+
}
|
| 276 |
+
],
|
| 277 |
+
"in_index": true
|
| 278 |
+
},
|
| 279 |
+
"31": {
|
| 280 |
+
"raw": 0.2292,
|
| 281 |
+
"skill": 0.1605,
|
| 282 |
+
"coverage": 1.0,
|
| 283 |
+
"pending": 0.0,
|
| 284 |
+
"random": 0.0819,
|
| 285 |
+
"tracks": [],
|
| 286 |
+
"in_index": true
|
| 287 |
+
},
|
| 288 |
+
"36": {
|
| 289 |
+
"raw": 0.177,
|
| 290 |
+
"skill": 0.1387,
|
| 291 |
+
"coverage": 1.0,
|
| 292 |
+
"pending": 0.0,
|
| 293 |
+
"random": 0.0444,
|
| 294 |
+
"tracks": [],
|
| 295 |
+
"in_index": true
|
| 296 |
+
},
|
| 297 |
+
"37": {
|
| 298 |
+
"raw": 0.5199,
|
| 299 |
+
"skill": 0.3978,
|
| 300 |
+
"coverage": 1.0,
|
| 301 |
+
"pending": 0.0,
|
| 302 |
+
"random": 0.2027,
|
| 303 |
+
"tracks": [],
|
| 304 |
+
"in_index": true
|
| 305 |
+
},
|
| 306 |
+
"40": {
|
| 307 |
+
"raw": 0.5057,
|
| 308 |
+
"skill": 0.3641,
|
| 309 |
+
"coverage": 1.0,
|
| 310 |
+
"pending": 0.0,
|
| 311 |
+
"random": 0.2227,
|
| 312 |
+
"tracks": [
|
| 313 |
+
{
|
| 314 |
+
"track": "A \u00b7 Arabic",
|
| 315 |
+
"score": 0.3205,
|
| 316 |
+
"headline": false
|
| 317 |
+
},
|
| 318 |
+
{
|
| 319 |
+
"track": "A \u00b7 English",
|
| 320 |
+
"score": 0.5057,
|
| 321 |
+
"headline": true
|
| 322 |
+
},
|
| 323 |
+
{
|
| 324 |
+
"track": "C \u00b7 Arabic pairs",
|
| 325 |
+
"score": 0.79,
|
| 326 |
+
"headline": false
|
| 327 |
+
},
|
| 328 |
+
{
|
| 329 |
+
"track": "C \u00b7 English pairs",
|
| 330 |
+
"score": 0.95,
|
| 331 |
+
"headline": false
|
| 332 |
+
}
|
| 333 |
+
],
|
| 334 |
+
"in_index": true
|
| 335 |
+
},
|
| 336 |
+
"41": {
|
| 337 |
+
"raw": 0.7805,
|
| 338 |
+
"skill": 0.6707,
|
| 339 |
+
"coverage": 1.0,
|
| 340 |
+
"pending": 0.0,
|
| 341 |
+
"random": 0.3333,
|
| 342 |
+
"tracks": [],
|
| 343 |
+
"in_index": true
|
| 344 |
+
},
|
| 345 |
+
"43": {
|
| 346 |
+
"raw": 0.5474,
|
| 347 |
+
"skill": 0.2819,
|
| 348 |
+
"coverage": 1.0,
|
| 349 |
+
"pending": 0.0,
|
| 350 |
+
"random": 0.3697,
|
| 351 |
+
"tracks": [],
|
| 352 |
+
"in_index": true
|
| 353 |
+
},
|
| 354 |
+
"44": {
|
| 355 |
+
"raw": 0.6614,
|
| 356 |
+
"skill": 0.3228,
|
| 357 |
+
"coverage": 1.0,
|
| 358 |
+
"pending": 0.0,
|
| 359 |
+
"random": 0.5,
|
| 360 |
+
"tracks": [],
|
| 361 |
+
"in_index": true
|
| 362 |
+
},
|
| 363 |
+
"50": {
|
| 364 |
+
"raw": 0.4553,
|
| 365 |
+
"skill": 0.2094,
|
| 366 |
+
"coverage": 1.0,
|
| 367 |
+
"pending": 0.0,
|
| 368 |
+
"random": 0.311,
|
| 369 |
+
"tracks": [],
|
| 370 |
+
"in_index": true
|
| 371 |
+
}
|
| 372 |
+
},
|
| 373 |
+
"benchmarks": {
|
| 374 |
+
"1": {
|
| 375 |
+
"catalog_id": 1,
|
| 376 |
+
"dataset": "BFCL",
|
| 377 |
+
"requests": 1694,
|
| 378 |
+
"answered": 1694,
|
| 379 |
+
"unsupported": 0,
|
| 380 |
+
"errors": 0,
|
| 381 |
+
"abstained": 0,
|
| 382 |
+
"pending": 0,
|
| 383 |
+
"scored_requests": 1694,
|
| 384 |
+
"metric": "case exact accuracy",
|
| 385 |
+
"score": 0.8926,
|
| 386 |
+
"reference_same_cases": null,
|
| 387 |
+
"median_ms": 139.8,
|
| 388 |
+
"in_index": true
|
| 389 |
+
},
|
| 390 |
+
"2": {
|
| 391 |
+
"catalog_id": 2,
|
| 392 |
+
"dataset": "ToolRet",
|
| 393 |
+
"requests": 7704,
|
| 394 |
+
"answered": 7704,
|
| 395 |
+
"unsupported": 0,
|
| 396 |
+
"errors": 0,
|
| 397 |
+
"abstained": 0,
|
| 398 |
+
"pending": 0,
|
| 399 |
+
"scored_requests": 7704,
|
| 400 |
+
"metric": "nDCG@10",
|
| 401 |
+
"score": 0.4222,
|
| 402 |
+
"reference_same_cases": null,
|
| 403 |
+
"median_ms": 1192.7,
|
| 404 |
+
"in_index": true
|
| 405 |
+
},
|
| 406 |
+
"3": {
|
| 407 |
+
"catalog_id": 3,
|
| 408 |
+
"dataset": "API-Bank",
|
| 409 |
+
"requests": 508,
|
| 410 |
+
"answered": 508,
|
| 411 |
+
"unsupported": 0,
|
| 412 |
+
"errors": 0,
|
| 413 |
+
"abstained": 0,
|
| 414 |
+
"pending": 0,
|
| 415 |
+
"scored_requests": 508,
|
| 416 |
+
"metric": "accuracy",
|
| 417 |
+
"score": 0.8386,
|
| 418 |
+
"reference_same_cases": null,
|
| 419 |
+
"median_ms": 640.5,
|
| 420 |
+
"in_index": false
|
| 421 |
+
},
|
| 422 |
+
"4": {
|
| 423 |
+
"catalog_id": 4,
|
| 424 |
+
"dataset": "BANKING77",
|
| 425 |
+
"requests": 3080,
|
| 426 |
+
"answered": 3080,
|
| 427 |
+
"unsupported": 0,
|
| 428 |
+
"errors": 0,
|
| 429 |
+
"abstained": 0,
|
| 430 |
+
"pending": 0,
|
| 431 |
+
"scored_requests": 3080,
|
| 432 |
+
"metric": "macro-F1",
|
| 433 |
+
"score": 0.8177,
|
| 434 |
+
"reference_same_cases": null,
|
| 435 |
+
"median_ms": 103.1,
|
| 436 |
+
"in_index": false
|
| 437 |
+
},
|
| 438 |
+
"5": {
|
| 439 |
+
"catalog_id": 5,
|
| 440 |
+
"dataset": "CLINC150+OOS",
|
| 441 |
+
"requests": 5500,
|
| 442 |
+
"answered": 5500,
|
| 443 |
+
"unsupported": 0,
|
| 444 |
+
"errors": 0,
|
| 445 |
+
"abstained": 0,
|
| 446 |
+
"pending": 0,
|
| 447 |
+
"scored_requests": 5500,
|
| 448 |
+
"metric": "macro-F1",
|
| 449 |
+
"score": 0.8381,
|
| 450 |
+
"reference_same_cases": null,
|
| 451 |
+
"median_ms": 172.3,
|
| 452 |
+
"in_index": false
|
| 453 |
+
},
|
| 454 |
+
"6": {
|
| 455 |
+
"catalog_id": 6,
|
| 456 |
+
"dataset": "RouterBench",
|
| 457 |
+
"requests": 10000,
|
| 458 |
+
"answered": 10000,
|
| 459 |
+
"unsupported": 0,
|
| 460 |
+
"errors": 0,
|
| 461 |
+
"abstained": 0,
|
| 462 |
+
"pending": 0,
|
| 463 |
+
"scored_requests": 10000,
|
| 464 |
+
"metric": "selected quality (quality objective)",
|
| 465 |
+
"score": 0.7989,
|
| 466 |
+
"reference_same_cases": null,
|
| 467 |
+
"median_ms": 392.5,
|
| 468 |
+
"tracks": {
|
| 469 |
+
"RouterBench-0shot": {
|
| 470 |
+
"metric": "selected quality (quality objective)",
|
| 471 |
+
"score": 0.7863581850889466,
|
| 472 |
+
"scored_requests": 5003
|
| 473 |
+
},
|
| 474 |
+
"RouterBench-5shot": {
|
| 475 |
+
"metric": "selected quality (quality objective)",
|
| 476 |
+
"score": 0.8114368621172704,
|
| 477 |
+
"scored_requests": 4997
|
| 478 |
+
}
|
| 479 |
+
},
|
| 480 |
+
"in_index": true
|
| 481 |
+
},
|
| 482 |
+
"9": {
|
| 483 |
+
"catalog_id": 9,
|
| 484 |
+
"dataset": "Home appliance simulator",
|
| 485 |
+
"requests": 160,
|
| 486 |
+
"answered": 160,
|
| 487 |
+
"unsupported": 0,
|
| 488 |
+
"errors": 0,
|
| 489 |
+
"abstained": 0,
|
| 490 |
+
"pending": 0,
|
| 491 |
+
"scored_requests": 160,
|
| 492 |
+
"metric": "case exact accuracy",
|
| 493 |
+
"score": 0.3063,
|
| 494 |
+
"reference_same_cases": null,
|
| 495 |
+
"median_ms": 1713.5,
|
| 496 |
+
"in_index": false
|
| 497 |
+
},
|
| 498 |
+
"10": {
|
| 499 |
+
"catalog_id": 10,
|
| 500 |
+
"dataset": "SGD/SGD-X",
|
| 501 |
+
"requests": 2500,
|
| 502 |
+
"answered": 2500,
|
| 503 |
+
"unsupported": 0,
|
| 504 |
+
"errors": 0,
|
| 505 |
+
"abstained": 0,
|
| 506 |
+
"pending": 0,
|
| 507 |
+
"scored_requests": 2500,
|
| 508 |
+
"metric": "macro-F1",
|
| 509 |
+
"score": 0.1982,
|
| 510 |
+
"reference_same_cases": null,
|
| 511 |
+
"median_ms": 106.5,
|
| 512 |
+
"in_index": false
|
| 513 |
+
},
|
| 514 |
+
"11": {
|
| 515 |
+
"catalog_id": 11,
|
| 516 |
+
"dataset": "ContractNLI",
|
| 517 |
+
"requests": 123,
|
| 518 |
+
"answered": 123,
|
| 519 |
+
"unsupported": 0,
|
| 520 |
+
"errors": 0,
|
| 521 |
+
"abstained": 0,
|
| 522 |
+
"pending": 0,
|
| 523 |
+
"scored_requests": 123,
|
| 524 |
+
"metric": "macro-F1",
|
| 525 |
+
"score": 0.8171,
|
| 526 |
+
"reference_same_cases": null,
|
| 527 |
+
"median_ms": 2966.1,
|
| 528 |
+
"in_index": true
|
| 529 |
+
},
|
| 530 |
+
"12": {
|
| 531 |
+
"catalog_id": 12,
|
| 532 |
+
"dataset": "ANLI",
|
| 533 |
+
"requests": 3200,
|
| 534 |
+
"answered": 3200,
|
| 535 |
+
"unsupported": 0,
|
| 536 |
+
"errors": 0,
|
| 537 |
+
"abstained": 0,
|
| 538 |
+
"pending": 0,
|
| 539 |
+
"scored_requests": 3200,
|
| 540 |
+
"metric": "macro-F1",
|
| 541 |
+
"score": 0.6899,
|
| 542 |
+
"reference_same_cases": null,
|
| 543 |
+
"median_ms": 19.4,
|
| 544 |
+
"in_index": false
|
| 545 |
+
},
|
| 546 |
+
"20": {
|
| 547 |
+
"catalog_id": 20,
|
| 548 |
+
"dataset": "BPoMP",
|
| 549 |
+
"requests": 5000,
|
| 550 |
+
"answered": 5000,
|
| 551 |
+
"unsupported": 0,
|
| 552 |
+
"errors": 0,
|
| 553 |
+
"abstained": 0,
|
| 554 |
+
"pending": 0,
|
| 555 |
+
"scored_requests": 5000,
|
| 556 |
+
"metric": "accuracy",
|
| 557 |
+
"score": 0.8398,
|
| 558 |
+
"reference_same_cases": null,
|
| 559 |
+
"median_ms": 18.7,
|
| 560 |
+
"in_index": true
|
| 561 |
+
},
|
| 562 |
+
"21": {
|
| 563 |
+
"catalog_id": 21,
|
| 564 |
+
"dataset": "Humicroedit",
|
| 565 |
+
"requests": 2628,
|
| 566 |
+
"answered": 2628,
|
| 567 |
+
"unsupported": 0,
|
| 568 |
+
"errors": 0,
|
| 569 |
+
"abstained": 0,
|
| 570 |
+
"pending": 0,
|
| 571 |
+
"scored_requests": 2628,
|
| 572 |
+
"metric": "accuracy",
|
| 573 |
+
"score": 0.6381,
|
| 574 |
+
"reference_same_cases": null,
|
| 575 |
+
"median_ms": 11.8,
|
| 576 |
+
"in_index": true
|
| 577 |
+
},
|
| 578 |
+
"22": {
|
| 579 |
+
"catalog_id": 22,
|
| 580 |
+
"dataset": "POP909-CL",
|
| 581 |
+
"requests": 2000,
|
| 582 |
+
"answered": 2000,
|
| 583 |
+
"unsupported": 0,
|
| 584 |
+
"errors": 0,
|
| 585 |
+
"abstained": 0,
|
| 586 |
+
"pending": 0,
|
| 587 |
+
"scored_requests": 2000,
|
| 588 |
+
"metric": "accuracy",
|
| 589 |
+
"score": 0.0825,
|
| 590 |
+
"reference_same_cases": null,
|
| 591 |
+
"median_ms": 843.5,
|
| 592 |
+
"in_index": true
|
| 593 |
+
},
|
| 594 |
+
"23": {
|
| 595 |
+
"catalog_id": 23,
|
| 596 |
+
"dataset": "cfcolor",
|
| 597 |
+
"requests": 5000,
|
| 598 |
+
"answered": 5000,
|
| 599 |
+
"unsupported": 0,
|
| 600 |
+
"errors": 0,
|
| 601 |
+
"abstained": 0,
|
| 602 |
+
"pending": 0,
|
| 603 |
+
"scored_requests": 5000,
|
| 604 |
+
"metric": "accuracy",
|
| 605 |
+
"score": 0.5886,
|
| 606 |
+
"reference_same_cases": null,
|
| 607 |
+
"median_ms": 50.0,
|
| 608 |
+
"in_index": true
|
| 609 |
+
},
|
| 610 |
+
"24": {
|
| 611 |
+
"catalog_id": 24,
|
| 612 |
+
"dataset": "MMLU",
|
| 613 |
+
"requests": 14033,
|
| 614 |
+
"answered": 14033,
|
| 615 |
+
"unsupported": 0,
|
| 616 |
+
"errors": 0,
|
| 617 |
+
"abstained": 0,
|
| 618 |
+
"pending": 0,
|
| 619 |
+
"scored_requests": 14033,
|
| 620 |
+
"metric": "accuracy",
|
| 621 |
+
"score": 0.788,
|
| 622 |
+
"reference_same_cases": null,
|
| 623 |
+
"median_ms": 18.2,
|
| 624 |
+
"in_index": true
|
| 625 |
+
},
|
| 626 |
+
"25": {
|
| 627 |
+
"catalog_id": 25,
|
| 628 |
+
"dataset": "GPQA Diamond",
|
| 629 |
+
"requests": 196,
|
| 630 |
+
"answered": 196,
|
| 631 |
+
"unsupported": 0,
|
| 632 |
+
"errors": 0,
|
| 633 |
+
"abstained": 0,
|
| 634 |
+
"pending": 0,
|
| 635 |
+
"scored_requests": 196,
|
| 636 |
+
"metric": "accuracy",
|
| 637 |
+
"score": 0.4082,
|
| 638 |
+
"reference_same_cases": null,
|
| 639 |
+
"median_ms": 34.6,
|
| 640 |
+
"in_index": true
|
| 641 |
+
},
|
| 642 |
+
"26": {
|
| 643 |
+
"catalog_id": 26,
|
| 644 |
+
"dataset": "ARC-Easy",
|
| 645 |
+
"requests": 2376,
|
| 646 |
+
"answered": 2376,
|
| 647 |
+
"unsupported": 0,
|
| 648 |
+
"errors": 0,
|
| 649 |
+
"abstained": 0,
|
| 650 |
+
"pending": 0,
|
| 651 |
+
"scored_requests": 2376,
|
| 652 |
+
"metric": "accuracy",
|
| 653 |
+
"score": 0.9798,
|
| 654 |
+
"reference_same_cases": null,
|
| 655 |
+
"median_ms": 16.1,
|
| 656 |
+
"in_index": false
|
| 657 |
+
},
|
| 658 |
+
"27": {
|
| 659 |
+
"catalog_id": 27,
|
| 660 |
+
"dataset": "ARC-Challenge",
|
| 661 |
+
"requests": 1172,
|
| 662 |
+
"answered": 1172,
|
| 663 |
+
"unsupported": 0,
|
| 664 |
+
"errors": 0,
|
| 665 |
+
"abstained": 0,
|
| 666 |
+
"pending": 0,
|
| 667 |
+
"scored_requests": 1172,
|
| 668 |
+
"metric": "accuracy",
|
| 669 |
+
"score": 0.9514,
|
| 670 |
+
"reference_same_cases": null,
|
| 671 |
+
"median_ms": 16.9,
|
| 672 |
+
"in_index": false
|
| 673 |
+
},
|
| 674 |
+
"28": {
|
| 675 |
+
"catalog_id": 28,
|
| 676 |
+
"dataset": "WinoGrande",
|
| 677 |
+
"requests": 1267,
|
| 678 |
+
"answered": 1267,
|
| 679 |
+
"unsupported": 0,
|
| 680 |
+
"errors": 0,
|
| 681 |
+
"abstained": 0,
|
| 682 |
+
"pending": 0,
|
| 683 |
+
"scored_requests": 1267,
|
| 684 |
+
"metric": "accuracy",
|
| 685 |
+
"score": 0.7419,
|
| 686 |
+
"reference_same_cases": null,
|
| 687 |
+
"median_ms": 12.0,
|
| 688 |
+
"in_index": false
|
| 689 |
+
},
|
| 690 |
+
"29": {
|
| 691 |
+
"catalog_id": 29,
|
| 692 |
+
"dataset": "HellaSwag",
|
| 693 |
+
"requests": 10042,
|
| 694 |
+
"answered": 10042,
|
| 695 |
+
"unsupported": 0,
|
| 696 |
+
"errors": 0,
|
| 697 |
+
"abstained": 0,
|
| 698 |
+
"pending": 0,
|
| 699 |
+
"scored_requests": 10042,
|
| 700 |
+
"metric": "accuracy",
|
| 701 |
+
"score": 0.867,
|
| 702 |
+
"reference_same_cases": null,
|
| 703 |
+
"median_ms": 27.6,
|
| 704 |
+
"in_index": false
|
| 705 |
+
},
|
| 706 |
+
"30": {
|
| 707 |
+
"catalog_id": 30,
|
| 708 |
+
"dataset": "GSM8K",
|
| 709 |
+
"requests": 2638,
|
| 710 |
+
"answered": 2638,
|
| 711 |
+
"unsupported": 0,
|
| 712 |
+
"errors": 0,
|
| 713 |
+
"abstained": 0,
|
| 714 |
+
"pending": 0,
|
| 715 |
+
"scored_requests": 2638,
|
| 716 |
+
"metric": "accuracy",
|
| 717 |
+
"score": 0.6585,
|
| 718 |
+
"reference_same_cases": null,
|
| 719 |
+
"median_ms": 24.3,
|
| 720 |
+
"tracks": {
|
| 721 |
+
"GSM8K-10choice": {
|
| 722 |
+
"metric": "accuracy",
|
| 723 |
+
"score": 0.6072782410917361,
|
| 724 |
+
"scored_requests": 1319
|
| 725 |
+
},
|
| 726 |
+
"GSM8K-4choice": {
|
| 727 |
+
"metric": "accuracy",
|
| 728 |
+
"score": 0.709628506444276,
|
| 729 |
+
"scored_requests": 1319
|
| 730 |
+
}
|
| 731 |
+
},
|
| 732 |
+
"in_index": true
|
| 733 |
+
},
|
| 734 |
+
"31": {
|
| 735 |
+
"catalog_id": 31,
|
| 736 |
+
"dataset": "ChessBench",
|
| 737 |
+
"requests": 5000,
|
| 738 |
+
"answered": 5000,
|
| 739 |
+
"unsupported": 0,
|
| 740 |
+
"errors": 0,
|
| 741 |
+
"abstained": 0,
|
| 742 |
+
"pending": 0,
|
| 743 |
+
"scored_requests": 5000,
|
| 744 |
+
"metric": "accuracy",
|
| 745 |
+
"score": 0.2292,
|
| 746 |
+
"reference_same_cases": null,
|
| 747 |
+
"median_ms": 132.4,
|
| 748 |
+
"in_index": true
|
| 749 |
+
},
|
| 750 |
+
"32": {
|
| 751 |
+
"catalog_id": 32,
|
| 752 |
+
"dataset": "MuSR",
|
| 753 |
+
"requests": 752,
|
| 754 |
+
"answered": 752,
|
| 755 |
+
"unsupported": 0,
|
| 756 |
+
"errors": 0,
|
| 757 |
+
"abstained": 0,
|
| 758 |
+
"pending": 0,
|
| 759 |
+
"scored_requests": 752,
|
| 760 |
+
"metric": "accuracy",
|
| 761 |
+
"score": 0.5705,
|
| 762 |
+
"reference_same_cases": null,
|
| 763 |
+
"median_ms": 93.7,
|
| 764 |
+
"in_index": false
|
| 765 |
+
},
|
| 766 |
+
"33": {
|
| 767 |
+
"catalog_id": 33,
|
| 768 |
+
"dataset": "SATA-Bench",
|
| 769 |
+
"requests": 1650,
|
| 770 |
+
"answered": 1650,
|
| 771 |
+
"unsupported": 0,
|
| 772 |
+
"errors": 0,
|
| 773 |
+
"abstained": 0,
|
| 774 |
+
"pending": 0,
|
| 775 |
+
"scored_requests": 1650,
|
| 776 |
+
"metric": "case exact accuracy",
|
| 777 |
+
"score": 0.3467,
|
| 778 |
+
"reference_same_cases": null,
|
| 779 |
+
"median_ms": 321.2,
|
| 780 |
+
"in_index": false
|
| 781 |
+
},
|
| 782 |
+
"34": {
|
| 783 |
+
"catalog_id": 34,
|
| 784 |
+
"dataset": "SimpleBench",
|
| 785 |
+
"requests": 10,
|
| 786 |
+
"answered": 10,
|
| 787 |
+
"unsupported": 0,
|
| 788 |
+
"errors": 0,
|
| 789 |
+
"abstained": 0,
|
| 790 |
+
"pending": 0,
|
| 791 |
+
"scored_requests": 10,
|
| 792 |
+
"metric": "accuracy",
|
| 793 |
+
"score": 0.1,
|
| 794 |
+
"reference_same_cases": null,
|
| 795 |
+
"median_ms": 21.2,
|
| 796 |
+
"in_index": false
|
| 797 |
+
},
|
| 798 |
+
"36": {
|
| 799 |
+
"catalog_id": 36,
|
| 800 |
+
"dataset": "BRIGHT",
|
| 801 |
+
"requests": 1297,
|
| 802 |
+
"answered": 1297,
|
| 803 |
+
"unsupported": 0,
|
| 804 |
+
"errors": 0,
|
| 805 |
+
"abstained": 0,
|
| 806 |
+
"pending": 0,
|
| 807 |
+
"scored_requests": 1297,
|
| 808 |
+
"metric": "nDCG@10",
|
| 809 |
+
"score": 0.177,
|
| 810 |
+
"reference_same_cases": null,
|
| 811 |
+
"median_ms": 1529.6,
|
| 812 |
+
"in_index": true
|
| 813 |
+
},
|
| 814 |
+
"37": {
|
| 815 |
+
"catalog_id": 37,
|
| 816 |
+
"dataset": "Amazon ESCI",
|
| 817 |
+
"requests": 5000,
|
| 818 |
+
"answered": 5000,
|
| 819 |
+
"unsupported": 0,
|
| 820 |
+
"errors": 0,
|
| 821 |
+
"abstained": 0,
|
| 822 |
+
"pending": 0,
|
| 823 |
+
"scored_requests": 5000,
|
| 824 |
+
"metric": "macro-F1",
|
| 825 |
+
"score": 0.5199,
|
| 826 |
+
"reference_same_cases": null,
|
| 827 |
+
"median_ms": 39.0,
|
| 828 |
+
"in_index": true
|
| 829 |
+
},
|
| 830 |
+
"38": {
|
| 831 |
+
"catalog_id": 38,
|
| 832 |
+
"dataset": "ACOS",
|
| 833 |
+
"requests": 5479,
|
| 834 |
+
"answered": 5479,
|
| 835 |
+
"unsupported": 0,
|
| 836 |
+
"errors": 0,
|
| 837 |
+
"abstained": 0,
|
| 838 |
+
"pending": 0,
|
| 839 |
+
"scored_requests": 5479,
|
| 840 |
+
"metric": "case exact accuracy",
|
| 841 |
+
"score": 0.0579,
|
| 842 |
+
"reference_same_cases": null,
|
| 843 |
+
"median_ms": 978.4,
|
| 844 |
+
"in_index": false
|
| 845 |
+
},
|
| 846 |
+
"39": {
|
| 847 |
+
"catalog_id": 39,
|
| 848 |
+
"dataset": "FinEntity",
|
| 849 |
+
"requests": 979,
|
| 850 |
+
"answered": 979,
|
| 851 |
+
"unsupported": 0,
|
| 852 |
+
"errors": 0,
|
| 853 |
+
"abstained": 0,
|
| 854 |
+
"pending": 0,
|
| 855 |
+
"scored_requests": 979,
|
| 856 |
+
"metric": "macro-F1",
|
| 857 |
+
"score": 0.8246,
|
| 858 |
+
"reference_same_cases": null,
|
| 859 |
+
"median_ms": 36.0,
|
| 860 |
+
"in_index": false
|
| 861 |
+
},
|
| 862 |
+
"40": {
|
| 863 |
+
"catalog_id": 40,
|
| 864 |
+
"dataset": "iSarcasmEval",
|
| 865 |
+
"requests": 4600,
|
| 866 |
+
"answered": 4600,
|
| 867 |
+
"unsupported": 0,
|
| 868 |
+
"errors": 0,
|
| 869 |
+
"abstained": 0,
|
| 870 |
+
"pending": 0,
|
| 871 |
+
"scored_requests": 4600,
|
| 872 |
+
"metric": "Sarcasm F1 \u00b7 track A, English",
|
| 873 |
+
"score": 0.5057,
|
| 874 |
+
"reference_same_cases": null,
|
| 875 |
+
"median_ms": 12.9,
|
| 876 |
+
"tracks": [
|
| 877 |
+
{
|
| 878 |
+
"track": "A \u00b7 Arabic",
|
| 879 |
+
"score": 0.3205,
|
| 880 |
+
"headline": false
|
| 881 |
+
},
|
| 882 |
+
{
|
| 883 |
+
"track": "A \u00b7 English",
|
| 884 |
+
"score": 0.5057,
|
| 885 |
+
"headline": true
|
| 886 |
+
},
|
| 887 |
+
{
|
| 888 |
+
"track": "C \u00b7 Arabic pairs",
|
| 889 |
+
"score": 0.79,
|
| 890 |
+
"headline": false
|
| 891 |
+
},
|
| 892 |
+
{
|
| 893 |
+
"track": "C \u00b7 English pairs",
|
| 894 |
+
"score": 0.95,
|
| 895 |
+
"headline": false
|
| 896 |
+
}
|
| 897 |
+
],
|
| 898 |
+
"in_index": true
|
| 899 |
+
},
|
| 900 |
+
"41": {
|
| 901 |
+
"catalog_id": 41,
|
| 902 |
+
"dataset": "VAST",
|
| 903 |
+
"requests": 3006,
|
| 904 |
+
"answered": 3006,
|
| 905 |
+
"unsupported": 0,
|
| 906 |
+
"errors": 0,
|
| 907 |
+
"abstained": 0,
|
| 908 |
+
"pending": 0,
|
| 909 |
+
"scored_requests": 3006,
|
| 910 |
+
"metric": "macro-F1",
|
| 911 |
+
"score": 0.7805,
|
| 912 |
+
"reference_same_cases": null,
|
| 913 |
+
"median_ms": 23.6,
|
| 914 |
+
"in_index": true
|
| 915 |
+
},
|
| 916 |
+
"42": {
|
| 917 |
+
"catalog_id": 42,
|
| 918 |
+
"dataset": "NLI4CT",
|
| 919 |
+
"requests": 5500,
|
| 920 |
+
"answered": 5500,
|
| 921 |
+
"unsupported": 0,
|
| 922 |
+
"errors": 0,
|
| 923 |
+
"abstained": 0,
|
| 924 |
+
"pending": 0,
|
| 925 |
+
"scored_requests": 5500,
|
| 926 |
+
"metric": "macro-F1",
|
| 927 |
+
"score": 0.8038,
|
| 928 |
+
"reference_same_cases": null,
|
| 929 |
+
"median_ms": 54.9,
|
| 930 |
+
"in_index": false
|
| 931 |
+
},
|
| 932 |
+
"43": {
|
| 933 |
+
"catalog_id": 43,
|
| 934 |
+
"dataset": "CRUXEval",
|
| 935 |
+
"requests": 570,
|
| 936 |
+
"answered": 570,
|
| 937 |
+
"unsupported": 0,
|
| 938 |
+
"errors": 0,
|
| 939 |
+
"abstained": 0,
|
| 940 |
+
"pending": 0,
|
| 941 |
+
"scored_requests": 570,
|
| 942 |
+
"metric": "accuracy",
|
| 943 |
+
"score": 0.5474,
|
| 944 |
+
"reference_same_cases": null,
|
| 945 |
+
"median_ms": 20.4,
|
| 946 |
+
"in_index": true
|
| 947 |
+
},
|
| 948 |
+
"44": {
|
| 949 |
+
"catalog_id": 44,
|
| 950 |
+
"dataset": "CLadder",
|
| 951 |
+
"requests": 5000,
|
| 952 |
+
"answered": 5000,
|
| 953 |
+
"unsupported": 0,
|
| 954 |
+
"errors": 0,
|
| 955 |
+
"abstained": 0,
|
| 956 |
+
"pending": 0,
|
| 957 |
+
"scored_requests": 5000,
|
| 958 |
+
"metric": "accuracy",
|
| 959 |
+
"score": 0.6614,
|
| 960 |
+
"reference_same_cases": null,
|
| 961 |
+
"median_ms": 19.9,
|
| 962 |
+
"in_index": true
|
| 963 |
+
},
|
| 964 |
+
"45": {
|
| 965 |
+
"catalog_id": 45,
|
| 966 |
+
"dataset": "HLE",
|
| 967 |
+
"requests": 501,
|
| 968 |
+
"answered": 501,
|
| 969 |
+
"unsupported": 0,
|
| 970 |
+
"errors": 0,
|
| 971 |
+
"abstained": 0,
|
| 972 |
+
"pending": 0,
|
| 973 |
+
"scored_requests": 501,
|
| 974 |
+
"metric": "accuracy",
|
| 975 |
+
"score": 0.1098,
|
| 976 |
+
"reference_same_cases": null,
|
| 977 |
+
"median_ms": 30.4,
|
| 978 |
+
"in_index": false
|
| 979 |
+
},
|
| 980 |
+
"48": {
|
| 981 |
+
"catalog_id": 48,
|
| 982 |
+
"dataset": "ForecastBench",
|
| 983 |
+
"requests": 10139,
|
| 984 |
+
"answered": 10139,
|
| 985 |
+
"unsupported": 0,
|
| 986 |
+
"errors": 0,
|
| 987 |
+
"abstained": 0,
|
| 988 |
+
"pending": 0,
|
| 989 |
+
"scored_requests": 10139,
|
| 990 |
+
"metric": "Brier (lower is better)",
|
| 991 |
+
"score": 0.1795,
|
| 992 |
+
"reference_same_cases": null,
|
| 993 |
+
"median_ms": 73.2,
|
| 994 |
+
"in_index": false
|
| 995 |
+
},
|
| 996 |
+
"50": {
|
| 997 |
+
"catalog_id": 50,
|
| 998 |
+
"dataset": "Habermas Machine",
|
| 999 |
+
"requests": 1676,
|
| 1000 |
+
"answered": 1676,
|
| 1001 |
+
"unsupported": 0,
|
| 1002 |
+
"errors": 0,
|
| 1003 |
+
"abstained": 0,
|
| 1004 |
+
"pending": 0,
|
| 1005 |
+
"scored_requests": 1676,
|
| 1006 |
+
"metric": "accuracy",
|
| 1007 |
+
"score": 0.4553,
|
| 1008 |
+
"reference_same_cases": null,
|
| 1009 |
+
"median_ms": 81.4,
|
| 1010 |
+
"in_index": true
|
| 1011 |
+
}
|
| 1012 |
+
},
|
| 1013 |
+
"frozen_panel": {
|
| 1014 |
+
"scores": {
|
| 1015 |
+
"balanced_raw": 43.85,
|
| 1016 |
+
"balanced_skill": 31.63,
|
| 1017 |
+
"breadth_skill": 27.36
|
| 1018 |
+
},
|
| 1019 |
+
"categories": [
|
| 1020 |
+
{
|
| 1021 |
+
"id": "knowledge",
|
| 1022 |
+
"raw": 0.6155,
|
| 1023 |
+
"skill": 0.4279,
|
| 1024 |
+
"coverage": 1.0,
|
| 1025 |
+
"pending": 0.0
|
| 1026 |
+
},
|
| 1027 |
+
{
|
| 1028 |
+
"id": "language",
|
| 1029 |
+
"raw": 0.5872,
|
| 1030 |
+
"skill": 0.4871,
|
| 1031 |
+
"coverage": 1.0,
|
| 1032 |
+
"pending": 0.0
|
| 1033 |
+
},
|
| 1034 |
+
{
|
| 1035 |
+
"id": "tools",
|
| 1036 |
+
"raw": 0.4227,
|
| 1037 |
+
"skill": 0.3486,
|
| 1038 |
+
"coverage": 0.6,
|
| 1039 |
+
"pending": 0.4
|
| 1040 |
+
},
|
| 1041 |
+
{
|
| 1042 |
+
"id": "games",
|
| 1043 |
+
"raw": 0.0458,
|
| 1044 |
+
"skill": 0.0321,
|
| 1045 |
+
"coverage": 0.2,
|
| 1046 |
+
"pending": 0.8
|
| 1047 |
+
},
|
| 1048 |
+
{
|
| 1049 |
+
"id": "arts",
|
| 1050 |
+
"raw": 0.5214,
|
| 1051 |
+
"skill": 0.2859,
|
| 1052 |
+
"coverage": 1.0,
|
| 1053 |
+
"pending": 0.0
|
| 1054 |
+
}
|
| 1055 |
+
],
|
| 1056 |
+
"coverage": 0.76,
|
| 1057 |
+
"pending": 0.24,
|
| 1058 |
+
"note": "25-benchmark frozen panel with the six interactive benchmarks scored as zero (provisional lower bound); not the headline index."
|
| 1059 |
+
},
|
| 1060 |
+
"formulas": [
|
| 1061 |
+
{
|
| 1062 |
+
"id": "balanced_raw",
|
| 1063 |
+
"normalization": "raw",
|
| 1064 |
+
"formula": "100*sum_c(0.2*category[c])"
|
| 1065 |
+
},
|
| 1066 |
+
{
|
| 1067 |
+
"id": "balanced_skill",
|
| 1068 |
+
"normalization": "skill",
|
| 1069 |
+
"formula": "100*sum_c(0.2*category[c])"
|
| 1070 |
+
},
|
| 1071 |
+
{
|
| 1072 |
+
"id": "breadth_skill",
|
| 1073 |
+
"normalization": "skill",
|
| 1074 |
+
"shift": 0.1,
|
| 1075 |
+
"formula": "100*(product_c((0.1+0.9*category[c])**0.2)-0.1)/0.9"
|
| 1076 |
+
}
|
| 1077 |
+
],
|
| 1078 |
+
"panel_id": "core25-observed-protocol-v1",
|
| 1079 |
+
"edition": "Decision Index 0.1 (archived 19-benchmark edition)"
|
| 1080 |
+
}
|
eval/decision-index-0.1-minus-dis.json
ADDED
|
@@ -0,0 +1,1070 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"engine": "blink-mimo-9b",
|
| 3 |
+
"generated_utc": "2026-09-24T14:48:25+00:00",
|
| 4 |
+
"suite": {
|
| 5 |
+
"requests": 132422,
|
| 6 |
+
"scoreable": 131980,
|
| 7 |
+
"excluded": 442,
|
| 8 |
+
"benchmarks": 37,
|
| 9 |
+
"rows_sha256": "750d353a3a83af615c67cfe9752e005bf09e6c28d9c4ba28d3a9f57ba8536cfd"
|
| 10 |
+
},
|
| 11 |
+
"completed": 132422,
|
| 12 |
+
"complete": true,
|
| 13 |
+
"counts": {
|
| 14 |
+
"ok": 132422
|
| 15 |
+
},
|
| 16 |
+
"latency_ms": {
|
| 17 |
+
"median": 49.9,
|
| 18 |
+
"p95": 1129.5,
|
| 19 |
+
"mean": 251.2
|
| 20 |
+
},
|
| 21 |
+
"decision_index": 56.6,
|
| 22 |
+
"scores": {
|
| 23 |
+
"balanced_raw": 56.6,
|
| 24 |
+
"balanced_skill": 42.28,
|
| 25 |
+
"breadth_skill": 40.47
|
| 26 |
+
},
|
| 27 |
+
"areas": [
|
| 28 |
+
{
|
| 29 |
+
"id": "knowledge",
|
| 30 |
+
"label": "Knowledge & Reasoning",
|
| 31 |
+
"raw": 0.5507,
|
| 32 |
+
"skill": 0.3832,
|
| 33 |
+
"coverage": 1.0,
|
| 34 |
+
"pending": 0.0,
|
| 35 |
+
"n": 6,
|
| 36 |
+
"benchmarks": [
|
| 37 |
+
24,
|
| 38 |
+
25,
|
| 39 |
+
30,
|
| 40 |
+
43,
|
| 41 |
+
44,
|
| 42 |
+
31
|
| 43 |
+
]
|
| 44 |
+
},
|
| 45 |
+
{
|
| 46 |
+
"id": "language",
|
| 47 |
+
"label": "Language Understanding",
|
| 48 |
+
"raw": 0.7036,
|
| 49 |
+
"skill": 0.5935,
|
| 50 |
+
"coverage": 1.0,
|
| 51 |
+
"pending": 0.0,
|
| 52 |
+
"n": 3,
|
| 53 |
+
"benchmarks": [
|
| 54 |
+
11,
|
| 55 |
+
40,
|
| 56 |
+
41
|
| 57 |
+
]
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"id": "retrieval",
|
| 61 |
+
"label": "Retrieval & Classification",
|
| 62 |
+
"raw": 0.3487,
|
| 63 |
+
"skill": 0.2687,
|
| 64 |
+
"coverage": 1.0,
|
| 65 |
+
"pending": 0.0,
|
| 66 |
+
"n": 2,
|
| 67 |
+
"benchmarks": [
|
| 68 |
+
36,
|
| 69 |
+
37
|
| 70 |
+
]
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"id": "tools",
|
| 74 |
+
"label": "Tools & Automation",
|
| 75 |
+
"raw": 0.7045,
|
| 76 |
+
"skill": 0.5811,
|
| 77 |
+
"coverage": 1.0,
|
| 78 |
+
"pending": 0.0,
|
| 79 |
+
"n": 3,
|
| 80 |
+
"benchmarks": [
|
| 81 |
+
1,
|
| 82 |
+
2,
|
| 83 |
+
6
|
| 84 |
+
]
|
| 85 |
+
},
|
| 86 |
+
{
|
| 87 |
+
"id": "arts",
|
| 88 |
+
"label": "Arts & Human Judgment",
|
| 89 |
+
"raw": 0.5225,
|
| 90 |
+
"skill": 0.2876,
|
| 91 |
+
"coverage": 1.0,
|
| 92 |
+
"pending": 0.0,
|
| 93 |
+
"n": 5,
|
| 94 |
+
"benchmarks": [
|
| 95 |
+
20,
|
| 96 |
+
21,
|
| 97 |
+
22,
|
| 98 |
+
23,
|
| 99 |
+
50
|
| 100 |
+
]
|
| 101 |
+
}
|
| 102 |
+
],
|
| 103 |
+
"index_benchmarks": {
|
| 104 |
+
"1": {
|
| 105 |
+
"raw": 0.8924,
|
| 106 |
+
"skill": 0.8547,
|
| 107 |
+
"coverage": 1.0,
|
| 108 |
+
"pending": 0.0,
|
| 109 |
+
"random": 0.2594,
|
| 110 |
+
"tracks": [],
|
| 111 |
+
"in_index": true
|
| 112 |
+
},
|
| 113 |
+
"2": {
|
| 114 |
+
"raw": 0.4219,
|
| 115 |
+
"skill": 0.3625,
|
| 116 |
+
"coverage": 1.0,
|
| 117 |
+
"pending": 0.0,
|
| 118 |
+
"random": 0.0933,
|
| 119 |
+
"tracks": [],
|
| 120 |
+
"in_index": true
|
| 121 |
+
},
|
| 122 |
+
"6": {
|
| 123 |
+
"raw": 0.7992,
|
| 124 |
+
"skill": 0.526,
|
| 125 |
+
"coverage": 1.0,
|
| 126 |
+
"pending": 0.0,
|
| 127 |
+
"random": 0.5248,
|
| 128 |
+
"tracks": [
|
| 129 |
+
{
|
| 130 |
+
"track": "RouterBench-0shot",
|
| 131 |
+
"score": 0.7868,
|
| 132 |
+
"headline": false
|
| 133 |
+
},
|
| 134 |
+
{
|
| 135 |
+
"track": "RouterBench-5shot",
|
| 136 |
+
"score": 0.8116,
|
| 137 |
+
"headline": false
|
| 138 |
+
}
|
| 139 |
+
],
|
| 140 |
+
"in_index": true
|
| 141 |
+
},
|
| 142 |
+
"7": {
|
| 143 |
+
"raw": 0.0,
|
| 144 |
+
"skill": 0.0,
|
| 145 |
+
"coverage": 0.0,
|
| 146 |
+
"pending": 1.0,
|
| 147 |
+
"random": null,
|
| 148 |
+
"tracks": [],
|
| 149 |
+
"in_index": false
|
| 150 |
+
},
|
| 151 |
+
"11": {
|
| 152 |
+
"raw": 0.8235,
|
| 153 |
+
"skill": 0.7448,
|
| 154 |
+
"coverage": 1.0,
|
| 155 |
+
"pending": 0.0,
|
| 156 |
+
"random": 0.3085,
|
| 157 |
+
"tracks": [],
|
| 158 |
+
"in_index": true
|
| 159 |
+
},
|
| 160 |
+
"13": {
|
| 161 |
+
"raw": 0.0,
|
| 162 |
+
"skill": 0.0,
|
| 163 |
+
"coverage": 0.0,
|
| 164 |
+
"pending": 1.0,
|
| 165 |
+
"random": null,
|
| 166 |
+
"tracks": [],
|
| 167 |
+
"in_index": false
|
| 168 |
+
},
|
| 169 |
+
"15": {
|
| 170 |
+
"raw": 0.0,
|
| 171 |
+
"skill": 0.0,
|
| 172 |
+
"coverage": 0.0,
|
| 173 |
+
"pending": 1.0,
|
| 174 |
+
"random": null,
|
| 175 |
+
"tracks": [],
|
| 176 |
+
"in_index": false
|
| 177 |
+
},
|
| 178 |
+
"16": {
|
| 179 |
+
"raw": 0.0,
|
| 180 |
+
"skill": 0.0,
|
| 181 |
+
"coverage": 0.0,
|
| 182 |
+
"pending": 1.0,
|
| 183 |
+
"random": null,
|
| 184 |
+
"tracks": [],
|
| 185 |
+
"in_index": false
|
| 186 |
+
},
|
| 187 |
+
"17": {
|
| 188 |
+
"raw": 0.0,
|
| 189 |
+
"skill": 0.0,
|
| 190 |
+
"coverage": 0.0,
|
| 191 |
+
"pending": 1.0,
|
| 192 |
+
"random": null,
|
| 193 |
+
"tracks": [],
|
| 194 |
+
"in_index": false
|
| 195 |
+
},
|
| 196 |
+
"19": {
|
| 197 |
+
"raw": 0.0,
|
| 198 |
+
"skill": 0.0,
|
| 199 |
+
"coverage": 0.0,
|
| 200 |
+
"pending": 1.0,
|
| 201 |
+
"random": null,
|
| 202 |
+
"tracks": [],
|
| 203 |
+
"in_index": false
|
| 204 |
+
},
|
| 205 |
+
"20": {
|
| 206 |
+
"raw": 0.8424,
|
| 207 |
+
"skill": 0.6848,
|
| 208 |
+
"coverage": 1.0,
|
| 209 |
+
"pending": 0.0,
|
| 210 |
+
"random": 0.5,
|
| 211 |
+
"tracks": [],
|
| 212 |
+
"in_index": true
|
| 213 |
+
},
|
| 214 |
+
"21": {
|
| 215 |
+
"raw": 0.6376,
|
| 216 |
+
"skill": 0.2753,
|
| 217 |
+
"coverage": 1.0,
|
| 218 |
+
"pending": 0.0,
|
| 219 |
+
"random": 0.5,
|
| 220 |
+
"tracks": [],
|
| 221 |
+
"in_index": true
|
| 222 |
+
},
|
| 223 |
+
"22": {
|
| 224 |
+
"raw": 0.076,
|
| 225 |
+
"skill": 0.0688,
|
| 226 |
+
"coverage": 1.0,
|
| 227 |
+
"pending": 0.0,
|
| 228 |
+
"random": 0.0078,
|
| 229 |
+
"tracks": [],
|
| 230 |
+
"in_index": true
|
| 231 |
+
},
|
| 232 |
+
"23": {
|
| 233 |
+
"raw": 0.5966,
|
| 234 |
+
"skill": 0.1932,
|
| 235 |
+
"coverage": 1.0,
|
| 236 |
+
"pending": 0.0,
|
| 237 |
+
"random": 0.5,
|
| 238 |
+
"tracks": [],
|
| 239 |
+
"in_index": true
|
| 240 |
+
},
|
| 241 |
+
"24": {
|
| 242 |
+
"raw": 0.8016,
|
| 243 |
+
"skill": 0.7355,
|
| 244 |
+
"coverage": 1.0,
|
| 245 |
+
"pending": 0.0,
|
| 246 |
+
"random": 0.25,
|
| 247 |
+
"tracks": [],
|
| 248 |
+
"in_index": true
|
| 249 |
+
},
|
| 250 |
+
"25": {
|
| 251 |
+
"raw": 0.3953,
|
| 252 |
+
"skill": 0.1938,
|
| 253 |
+
"coverage": 1.0,
|
| 254 |
+
"pending": 0.0,
|
| 255 |
+
"random": 0.25,
|
| 256 |
+
"tracks": [],
|
| 257 |
+
"in_index": true
|
| 258 |
+
},
|
| 259 |
+
"30": {
|
| 260 |
+
"raw": 0.6562,
|
| 261 |
+
"skill": 0.5854,
|
| 262 |
+
"coverage": 1.0,
|
| 263 |
+
"pending": 0.0,
|
| 264 |
+
"random": 0.25,
|
| 265 |
+
"tracks": [
|
| 266 |
+
{
|
| 267 |
+
"track": "GSM8K-4choice",
|
| 268 |
+
"score": 0.7069,
|
| 269 |
+
"headline": false
|
| 270 |
+
},
|
| 271 |
+
{
|
| 272 |
+
"track": "GSM8K-10choice",
|
| 273 |
+
"score": 0.6054,
|
| 274 |
+
"headline": false
|
| 275 |
+
}
|
| 276 |
+
],
|
| 277 |
+
"in_index": true
|
| 278 |
+
},
|
| 279 |
+
"31": {
|
| 280 |
+
"raw": 0.2291,
|
| 281 |
+
"skill": 0.1601,
|
| 282 |
+
"coverage": 1.0,
|
| 283 |
+
"pending": 0.0,
|
| 284 |
+
"random": 0.0821,
|
| 285 |
+
"tracks": [],
|
| 286 |
+
"in_index": true
|
| 287 |
+
},
|
| 288 |
+
"36": {
|
| 289 |
+
"raw": 0.1773,
|
| 290 |
+
"skill": 0.1394,
|
| 291 |
+
"coverage": 1.0,
|
| 292 |
+
"pending": 0.0,
|
| 293 |
+
"random": 0.044,
|
| 294 |
+
"tracks": [],
|
| 295 |
+
"in_index": true
|
| 296 |
+
},
|
| 297 |
+
"37": {
|
| 298 |
+
"raw": 0.5201,
|
| 299 |
+
"skill": 0.398,
|
| 300 |
+
"coverage": 1.0,
|
| 301 |
+
"pending": 0.0,
|
| 302 |
+
"random": 0.2027,
|
| 303 |
+
"tracks": [],
|
| 304 |
+
"in_index": true
|
| 305 |
+
},
|
| 306 |
+
"40": {
|
| 307 |
+
"raw": 0.5081,
|
| 308 |
+
"skill": 0.3672,
|
| 309 |
+
"coverage": 1.0,
|
| 310 |
+
"pending": 0.0,
|
| 311 |
+
"random": 0.2227,
|
| 312 |
+
"tracks": [
|
| 313 |
+
{
|
| 314 |
+
"track": "A \u00b7 Arabic",
|
| 315 |
+
"score": 0.3183,
|
| 316 |
+
"headline": false
|
| 317 |
+
},
|
| 318 |
+
{
|
| 319 |
+
"track": "A \u00b7 English",
|
| 320 |
+
"score": 0.5081,
|
| 321 |
+
"headline": true
|
| 322 |
+
},
|
| 323 |
+
{
|
| 324 |
+
"track": "C \u00b7 Arabic pairs",
|
| 325 |
+
"score": 0.7889,
|
| 326 |
+
"headline": false
|
| 327 |
+
},
|
| 328 |
+
{
|
| 329 |
+
"track": "C \u00b7 English pairs",
|
| 330 |
+
"score": 0.9497,
|
| 331 |
+
"headline": false
|
| 332 |
+
}
|
| 333 |
+
],
|
| 334 |
+
"in_index": true
|
| 335 |
+
},
|
| 336 |
+
"41": {
|
| 337 |
+
"raw": 0.7791,
|
| 338 |
+
"skill": 0.6686,
|
| 339 |
+
"coverage": 1.0,
|
| 340 |
+
"pending": 0.0,
|
| 341 |
+
"random": 0.3333,
|
| 342 |
+
"tracks": [],
|
| 343 |
+
"in_index": true
|
| 344 |
+
},
|
| 345 |
+
"43": {
|
| 346 |
+
"raw": 0.5606,
|
| 347 |
+
"skill": 0.3013,
|
| 348 |
+
"coverage": 1.0,
|
| 349 |
+
"pending": 0.0,
|
| 350 |
+
"random": 0.3711,
|
| 351 |
+
"tracks": [],
|
| 352 |
+
"in_index": true
|
| 353 |
+
},
|
| 354 |
+
"44": {
|
| 355 |
+
"raw": 0.6615,
|
| 356 |
+
"skill": 0.3229,
|
| 357 |
+
"coverage": 1.0,
|
| 358 |
+
"pending": 0.0,
|
| 359 |
+
"random": 0.5,
|
| 360 |
+
"tracks": [],
|
| 361 |
+
"in_index": true
|
| 362 |
+
},
|
| 363 |
+
"50": {
|
| 364 |
+
"raw": 0.4599,
|
| 365 |
+
"skill": 0.2161,
|
| 366 |
+
"coverage": 1.0,
|
| 367 |
+
"pending": 0.0,
|
| 368 |
+
"random": 0.311,
|
| 369 |
+
"tracks": [],
|
| 370 |
+
"in_index": true
|
| 371 |
+
}
|
| 372 |
+
},
|
| 373 |
+
"benchmarks": {
|
| 374 |
+
"1": {
|
| 375 |
+
"catalog_id": 1,
|
| 376 |
+
"dataset": "BFCL",
|
| 377 |
+
"requests": 1626,
|
| 378 |
+
"answered": 1626,
|
| 379 |
+
"unsupported": 0,
|
| 380 |
+
"errors": 0,
|
| 381 |
+
"abstained": 0,
|
| 382 |
+
"pending": 0,
|
| 383 |
+
"scored_requests": 1626,
|
| 384 |
+
"metric": "case exact accuracy",
|
| 385 |
+
"score": 0.8924,
|
| 386 |
+
"reference_same_cases": null,
|
| 387 |
+
"median_ms": 138.9,
|
| 388 |
+
"in_index": true
|
| 389 |
+
},
|
| 390 |
+
"2": {
|
| 391 |
+
"catalog_id": 2,
|
| 392 |
+
"dataset": "ToolRet",
|
| 393 |
+
"requests": 7636,
|
| 394 |
+
"answered": 7636,
|
| 395 |
+
"unsupported": 0,
|
| 396 |
+
"errors": 0,
|
| 397 |
+
"abstained": 0,
|
| 398 |
+
"pending": 0,
|
| 399 |
+
"scored_requests": 7636,
|
| 400 |
+
"metric": "nDCG@10",
|
| 401 |
+
"score": 0.4219,
|
| 402 |
+
"reference_same_cases": null,
|
| 403 |
+
"median_ms": 1193.7,
|
| 404 |
+
"in_index": true
|
| 405 |
+
},
|
| 406 |
+
"3": {
|
| 407 |
+
"catalog_id": 3,
|
| 408 |
+
"dataset": "API-Bank",
|
| 409 |
+
"requests": 440,
|
| 410 |
+
"answered": 440,
|
| 411 |
+
"unsupported": 0,
|
| 412 |
+
"errors": 0,
|
| 413 |
+
"abstained": 0,
|
| 414 |
+
"pending": 0,
|
| 415 |
+
"scored_requests": 440,
|
| 416 |
+
"metric": "accuracy",
|
| 417 |
+
"score": 0.8432,
|
| 418 |
+
"reference_same_cases": null,
|
| 419 |
+
"median_ms": 640.3,
|
| 420 |
+
"in_index": false
|
| 421 |
+
},
|
| 422 |
+
"4": {
|
| 423 |
+
"catalog_id": 4,
|
| 424 |
+
"dataset": "BANKING77",
|
| 425 |
+
"requests": 3012,
|
| 426 |
+
"answered": 3012,
|
| 427 |
+
"unsupported": 0,
|
| 428 |
+
"errors": 0,
|
| 429 |
+
"abstained": 0,
|
| 430 |
+
"pending": 0,
|
| 431 |
+
"scored_requests": 3012,
|
| 432 |
+
"metric": "macro-F1",
|
| 433 |
+
"score": 0.8162,
|
| 434 |
+
"reference_same_cases": null,
|
| 435 |
+
"median_ms": 103.1,
|
| 436 |
+
"in_index": false
|
| 437 |
+
},
|
| 438 |
+
"5": {
|
| 439 |
+
"catalog_id": 5,
|
| 440 |
+
"dataset": "CLINC150+OOS",
|
| 441 |
+
"requests": 5432,
|
| 442 |
+
"answered": 5432,
|
| 443 |
+
"unsupported": 0,
|
| 444 |
+
"errors": 0,
|
| 445 |
+
"abstained": 0,
|
| 446 |
+
"pending": 0,
|
| 447 |
+
"scored_requests": 5432,
|
| 448 |
+
"metric": "macro-F1",
|
| 449 |
+
"score": 0.8371,
|
| 450 |
+
"reference_same_cases": null,
|
| 451 |
+
"median_ms": 172.3,
|
| 452 |
+
"in_index": false
|
| 453 |
+
},
|
| 454 |
+
"6": {
|
| 455 |
+
"catalog_id": 6,
|
| 456 |
+
"dataset": "RouterBench",
|
| 457 |
+
"requests": 9864,
|
| 458 |
+
"answered": 9864,
|
| 459 |
+
"unsupported": 0,
|
| 460 |
+
"errors": 0,
|
| 461 |
+
"abstained": 0,
|
| 462 |
+
"pending": 0,
|
| 463 |
+
"scored_requests": 9864,
|
| 464 |
+
"metric": "selected quality (quality objective)",
|
| 465 |
+
"score": 0.7992,
|
| 466 |
+
"reference_same_cases": null,
|
| 467 |
+
"median_ms": 392.5,
|
| 468 |
+
"tracks": {
|
| 469 |
+
"RouterBench-0shot": {
|
| 470 |
+
"metric": "selected quality (quality objective)",
|
| 471 |
+
"score": 0.7868085106382978,
|
| 472 |
+
"scored_requests": 4935
|
| 473 |
+
},
|
| 474 |
+
"RouterBench-5shot": {
|
| 475 |
+
"metric": "selected quality (quality objective)",
|
| 476 |
+
"score": 0.8116250760803408,
|
| 477 |
+
"scored_requests": 4929
|
| 478 |
+
}
|
| 479 |
+
},
|
| 480 |
+
"in_index": true
|
| 481 |
+
},
|
| 482 |
+
"9": {
|
| 483 |
+
"catalog_id": 9,
|
| 484 |
+
"dataset": "Home appliance simulator",
|
| 485 |
+
"requests": 93,
|
| 486 |
+
"answered": 93,
|
| 487 |
+
"unsupported": 0,
|
| 488 |
+
"errors": 0,
|
| 489 |
+
"abstained": 0,
|
| 490 |
+
"pending": 0,
|
| 491 |
+
"scored_requests": 93,
|
| 492 |
+
"metric": "case exact accuracy",
|
| 493 |
+
"score": 0.2796,
|
| 494 |
+
"reference_same_cases": null,
|
| 495 |
+
"median_ms": 1712.7,
|
| 496 |
+
"in_index": false
|
| 497 |
+
},
|
| 498 |
+
"10": {
|
| 499 |
+
"catalog_id": 10,
|
| 500 |
+
"dataset": "SGD/SGD-X",
|
| 501 |
+
"requests": 2433,
|
| 502 |
+
"answered": 2433,
|
| 503 |
+
"unsupported": 0,
|
| 504 |
+
"errors": 0,
|
| 505 |
+
"abstained": 0,
|
| 506 |
+
"pending": 0,
|
| 507 |
+
"scored_requests": 2433,
|
| 508 |
+
"metric": "macro-F1",
|
| 509 |
+
"score": 0.1939,
|
| 510 |
+
"reference_same_cases": null,
|
| 511 |
+
"median_ms": 106.5,
|
| 512 |
+
"in_index": false
|
| 513 |
+
},
|
| 514 |
+
"11": {
|
| 515 |
+
"catalog_id": 11,
|
| 516 |
+
"dataset": "ContractNLI",
|
| 517 |
+
"requests": 56,
|
| 518 |
+
"answered": 56,
|
| 519 |
+
"unsupported": 0,
|
| 520 |
+
"errors": 0,
|
| 521 |
+
"abstained": 0,
|
| 522 |
+
"pending": 0,
|
| 523 |
+
"scored_requests": 56,
|
| 524 |
+
"metric": "macro-F1",
|
| 525 |
+
"score": 0.8235,
|
| 526 |
+
"reference_same_cases": null,
|
| 527 |
+
"median_ms": 2982.4,
|
| 528 |
+
"in_index": true
|
| 529 |
+
},
|
| 530 |
+
"12": {
|
| 531 |
+
"catalog_id": 12,
|
| 532 |
+
"dataset": "ANLI",
|
| 533 |
+
"requests": 3133,
|
| 534 |
+
"answered": 3133,
|
| 535 |
+
"unsupported": 0,
|
| 536 |
+
"errors": 0,
|
| 537 |
+
"abstained": 0,
|
| 538 |
+
"pending": 0,
|
| 539 |
+
"scored_requests": 3133,
|
| 540 |
+
"metric": "macro-F1",
|
| 541 |
+
"score": 0.6905,
|
| 542 |
+
"reference_same_cases": null,
|
| 543 |
+
"median_ms": 19.3,
|
| 544 |
+
"in_index": false
|
| 545 |
+
},
|
| 546 |
+
"20": {
|
| 547 |
+
"catalog_id": 20,
|
| 548 |
+
"dataset": "BPoMP",
|
| 549 |
+
"requests": 4696,
|
| 550 |
+
"answered": 4696,
|
| 551 |
+
"unsupported": 0,
|
| 552 |
+
"errors": 0,
|
| 553 |
+
"abstained": 0,
|
| 554 |
+
"pending": 0,
|
| 555 |
+
"scored_requests": 4696,
|
| 556 |
+
"metric": "accuracy",
|
| 557 |
+
"score": 0.842,
|
| 558 |
+
"reference_same_cases": null,
|
| 559 |
+
"median_ms": 18.7,
|
| 560 |
+
"in_index": true
|
| 561 |
+
},
|
| 562 |
+
"21": {
|
| 563 |
+
"catalog_id": 21,
|
| 564 |
+
"dataset": "Humicroedit",
|
| 565 |
+
"requests": 2561,
|
| 566 |
+
"answered": 2561,
|
| 567 |
+
"unsupported": 0,
|
| 568 |
+
"errors": 0,
|
| 569 |
+
"abstained": 0,
|
| 570 |
+
"pending": 0,
|
| 571 |
+
"scored_requests": 2561,
|
| 572 |
+
"metric": "accuracy",
|
| 573 |
+
"score": 0.6376,
|
| 574 |
+
"reference_same_cases": null,
|
| 575 |
+
"median_ms": 11.8,
|
| 576 |
+
"in_index": true
|
| 577 |
+
},
|
| 578 |
+
"22": {
|
| 579 |
+
"catalog_id": 22,
|
| 580 |
+
"dataset": "POP909-CL",
|
| 581 |
+
"requests": 1933,
|
| 582 |
+
"answered": 1933,
|
| 583 |
+
"unsupported": 0,
|
| 584 |
+
"errors": 0,
|
| 585 |
+
"abstained": 0,
|
| 586 |
+
"pending": 0,
|
| 587 |
+
"scored_requests": 1933,
|
| 588 |
+
"metric": "accuracy",
|
| 589 |
+
"score": 0.0823,
|
| 590 |
+
"reference_same_cases": null,
|
| 591 |
+
"median_ms": 842.8,
|
| 592 |
+
"in_index": true
|
| 593 |
+
},
|
| 594 |
+
"23": {
|
| 595 |
+
"catalog_id": 23,
|
| 596 |
+
"dataset": "cfcolor",
|
| 597 |
+
"requests": 4933,
|
| 598 |
+
"answered": 4933,
|
| 599 |
+
"unsupported": 0,
|
| 600 |
+
"errors": 0,
|
| 601 |
+
"abstained": 0,
|
| 602 |
+
"pending": 0,
|
| 603 |
+
"scored_requests": 4933,
|
| 604 |
+
"metric": "accuracy",
|
| 605 |
+
"score": 0.5883,
|
| 606 |
+
"reference_same_cases": null,
|
| 607 |
+
"median_ms": 50.0,
|
| 608 |
+
"in_index": true
|
| 609 |
+
},
|
| 610 |
+
"24": {
|
| 611 |
+
"catalog_id": 24,
|
| 612 |
+
"dataset": "MMLU",
|
| 613 |
+
"requests": 13966,
|
| 614 |
+
"answered": 13966,
|
| 615 |
+
"unsupported": 0,
|
| 616 |
+
"errors": 0,
|
| 617 |
+
"abstained": 0,
|
| 618 |
+
"pending": 0,
|
| 619 |
+
"scored_requests": 13966,
|
| 620 |
+
"metric": "accuracy",
|
| 621 |
+
"score": 0.7879,
|
| 622 |
+
"reference_same_cases": null,
|
| 623 |
+
"median_ms": 18.2,
|
| 624 |
+
"in_index": true
|
| 625 |
+
},
|
| 626 |
+
"25": {
|
| 627 |
+
"catalog_id": 25,
|
| 628 |
+
"dataset": "GPQA Diamond",
|
| 629 |
+
"requests": 129,
|
| 630 |
+
"answered": 129,
|
| 631 |
+
"unsupported": 0,
|
| 632 |
+
"errors": 0,
|
| 633 |
+
"abstained": 0,
|
| 634 |
+
"pending": 0,
|
| 635 |
+
"scored_requests": 129,
|
| 636 |
+
"metric": "accuracy",
|
| 637 |
+
"score": 0.3953,
|
| 638 |
+
"reference_same_cases": null,
|
| 639 |
+
"median_ms": 35.2,
|
| 640 |
+
"in_index": true
|
| 641 |
+
},
|
| 642 |
+
"26": {
|
| 643 |
+
"catalog_id": 26,
|
| 644 |
+
"dataset": "ARC-Easy",
|
| 645 |
+
"requests": 2309,
|
| 646 |
+
"answered": 2309,
|
| 647 |
+
"unsupported": 0,
|
| 648 |
+
"errors": 0,
|
| 649 |
+
"abstained": 0,
|
| 650 |
+
"pending": 0,
|
| 651 |
+
"scored_requests": 2309,
|
| 652 |
+
"metric": "accuracy",
|
| 653 |
+
"score": 0.9801,
|
| 654 |
+
"reference_same_cases": null,
|
| 655 |
+
"median_ms": 16.1,
|
| 656 |
+
"in_index": false
|
| 657 |
+
},
|
| 658 |
+
"27": {
|
| 659 |
+
"catalog_id": 27,
|
| 660 |
+
"dataset": "ARC-Challenge",
|
| 661 |
+
"requests": 1105,
|
| 662 |
+
"answered": 1105,
|
| 663 |
+
"unsupported": 0,
|
| 664 |
+
"errors": 0,
|
| 665 |
+
"abstained": 0,
|
| 666 |
+
"pending": 0,
|
| 667 |
+
"scored_requests": 1105,
|
| 668 |
+
"metric": "accuracy",
|
| 669 |
+
"score": 0.9502,
|
| 670 |
+
"reference_same_cases": null,
|
| 671 |
+
"median_ms": 16.9,
|
| 672 |
+
"in_index": false
|
| 673 |
+
},
|
| 674 |
+
"28": {
|
| 675 |
+
"catalog_id": 28,
|
| 676 |
+
"dataset": "WinoGrande",
|
| 677 |
+
"requests": 1200,
|
| 678 |
+
"answered": 1200,
|
| 679 |
+
"unsupported": 0,
|
| 680 |
+
"errors": 0,
|
| 681 |
+
"abstained": 0,
|
| 682 |
+
"pending": 0,
|
| 683 |
+
"scored_requests": 1200,
|
| 684 |
+
"metric": "accuracy",
|
| 685 |
+
"score": 0.74,
|
| 686 |
+
"reference_same_cases": null,
|
| 687 |
+
"median_ms": 12.0,
|
| 688 |
+
"in_index": false
|
| 689 |
+
},
|
| 690 |
+
"29": {
|
| 691 |
+
"catalog_id": 29,
|
| 692 |
+
"dataset": "HellaSwag",
|
| 693 |
+
"requests": 9975,
|
| 694 |
+
"answered": 9975,
|
| 695 |
+
"unsupported": 0,
|
| 696 |
+
"errors": 0,
|
| 697 |
+
"abstained": 0,
|
| 698 |
+
"pending": 0,
|
| 699 |
+
"scored_requests": 9975,
|
| 700 |
+
"metric": "accuracy",
|
| 701 |
+
"score": 0.8673,
|
| 702 |
+
"reference_same_cases": null,
|
| 703 |
+
"median_ms": 27.6,
|
| 704 |
+
"in_index": false
|
| 705 |
+
},
|
| 706 |
+
"30": {
|
| 707 |
+
"catalog_id": 30,
|
| 708 |
+
"dataset": "GSM8K",
|
| 709 |
+
"requests": 2504,
|
| 710 |
+
"answered": 2504,
|
| 711 |
+
"unsupported": 0,
|
| 712 |
+
"errors": 0,
|
| 713 |
+
"abstained": 0,
|
| 714 |
+
"pending": 0,
|
| 715 |
+
"scored_requests": 2504,
|
| 716 |
+
"metric": "accuracy",
|
| 717 |
+
"score": 0.6562,
|
| 718 |
+
"reference_same_cases": null,
|
| 719 |
+
"median_ms": 24.2,
|
| 720 |
+
"tracks": {
|
| 721 |
+
"GSM8K-10choice": {
|
| 722 |
+
"metric": "accuracy",
|
| 723 |
+
"score": 0.6054313099041534,
|
| 724 |
+
"scored_requests": 1252
|
| 725 |
+
},
|
| 726 |
+
"GSM8K-4choice": {
|
| 727 |
+
"metric": "accuracy",
|
| 728 |
+
"score": 0.7068690095846646,
|
| 729 |
+
"scored_requests": 1252
|
| 730 |
+
}
|
| 731 |
+
},
|
| 732 |
+
"in_index": true
|
| 733 |
+
},
|
| 734 |
+
"31": {
|
| 735 |
+
"catalog_id": 31,
|
| 736 |
+
"dataset": "ChessBench",
|
| 737 |
+
"requests": 4933,
|
| 738 |
+
"answered": 4933,
|
| 739 |
+
"unsupported": 0,
|
| 740 |
+
"errors": 0,
|
| 741 |
+
"abstained": 0,
|
| 742 |
+
"pending": 0,
|
| 743 |
+
"scored_requests": 4933,
|
| 744 |
+
"metric": "accuracy",
|
| 745 |
+
"score": 0.2291,
|
| 746 |
+
"reference_same_cases": null,
|
| 747 |
+
"median_ms": 132.4,
|
| 748 |
+
"in_index": true
|
| 749 |
+
},
|
| 750 |
+
"32": {
|
| 751 |
+
"catalog_id": 32,
|
| 752 |
+
"dataset": "MuSR",
|
| 753 |
+
"requests": 685,
|
| 754 |
+
"answered": 685,
|
| 755 |
+
"unsupported": 0,
|
| 756 |
+
"errors": 0,
|
| 757 |
+
"abstained": 0,
|
| 758 |
+
"pending": 0,
|
| 759 |
+
"scored_requests": 685,
|
| 760 |
+
"metric": "accuracy",
|
| 761 |
+
"score": 0.5635,
|
| 762 |
+
"reference_same_cases": null,
|
| 763 |
+
"median_ms": 93.4,
|
| 764 |
+
"in_index": false
|
| 765 |
+
},
|
| 766 |
+
"33": {
|
| 767 |
+
"catalog_id": 33,
|
| 768 |
+
"dataset": "SATA-Bench",
|
| 769 |
+
"requests": 1583,
|
| 770 |
+
"answered": 1583,
|
| 771 |
+
"unsupported": 0,
|
| 772 |
+
"errors": 0,
|
| 773 |
+
"abstained": 0,
|
| 774 |
+
"pending": 0,
|
| 775 |
+
"scored_requests": 1583,
|
| 776 |
+
"metric": "case exact accuracy",
|
| 777 |
+
"score": 0.3462,
|
| 778 |
+
"reference_same_cases": null,
|
| 779 |
+
"median_ms": 318.6,
|
| 780 |
+
"in_index": false
|
| 781 |
+
},
|
| 782 |
+
"36": {
|
| 783 |
+
"catalog_id": 36,
|
| 784 |
+
"dataset": "BRIGHT",
|
| 785 |
+
"requests": 1230,
|
| 786 |
+
"answered": 1230,
|
| 787 |
+
"unsupported": 0,
|
| 788 |
+
"errors": 0,
|
| 789 |
+
"abstained": 0,
|
| 790 |
+
"pending": 0,
|
| 791 |
+
"scored_requests": 1230,
|
| 792 |
+
"metric": "nDCG@10",
|
| 793 |
+
"score": 0.1773,
|
| 794 |
+
"reference_same_cases": null,
|
| 795 |
+
"median_ms": 1526.9,
|
| 796 |
+
"in_index": true
|
| 797 |
+
},
|
| 798 |
+
"37": {
|
| 799 |
+
"catalog_id": 37,
|
| 800 |
+
"dataset": "Amazon ESCI",
|
| 801 |
+
"requests": 4933,
|
| 802 |
+
"answered": 4933,
|
| 803 |
+
"unsupported": 0,
|
| 804 |
+
"errors": 0,
|
| 805 |
+
"abstained": 0,
|
| 806 |
+
"pending": 0,
|
| 807 |
+
"scored_requests": 4933,
|
| 808 |
+
"metric": "macro-F1",
|
| 809 |
+
"score": 0.5201,
|
| 810 |
+
"reference_same_cases": null,
|
| 811 |
+
"median_ms": 39.0,
|
| 812 |
+
"in_index": true
|
| 813 |
+
},
|
| 814 |
+
"38": {
|
| 815 |
+
"catalog_id": 38,
|
| 816 |
+
"dataset": "ACOS",
|
| 817 |
+
"requests": 5212,
|
| 818 |
+
"answered": 5212,
|
| 819 |
+
"unsupported": 0,
|
| 820 |
+
"errors": 0,
|
| 821 |
+
"abstained": 0,
|
| 822 |
+
"pending": 0,
|
| 823 |
+
"scored_requests": 5212,
|
| 824 |
+
"metric": "case exact accuracy",
|
| 825 |
+
"score": 0.0593,
|
| 826 |
+
"reference_same_cases": null,
|
| 827 |
+
"median_ms": 978.3,
|
| 828 |
+
"in_index": false
|
| 829 |
+
},
|
| 830 |
+
"39": {
|
| 831 |
+
"catalog_id": 39,
|
| 832 |
+
"dataset": "FinEntity",
|
| 833 |
+
"requests": 912,
|
| 834 |
+
"answered": 912,
|
| 835 |
+
"unsupported": 0,
|
| 836 |
+
"errors": 0,
|
| 837 |
+
"abstained": 0,
|
| 838 |
+
"pending": 0,
|
| 839 |
+
"scored_requests": 912,
|
| 840 |
+
"metric": "macro-F1",
|
| 841 |
+
"score": 0.8244,
|
| 842 |
+
"reference_same_cases": null,
|
| 843 |
+
"median_ms": 36.2,
|
| 844 |
+
"in_index": false
|
| 845 |
+
},
|
| 846 |
+
"40": {
|
| 847 |
+
"catalog_id": 40,
|
| 848 |
+
"dataset": "iSarcasmEval",
|
| 849 |
+
"requests": 4533,
|
| 850 |
+
"answered": 4533,
|
| 851 |
+
"unsupported": 0,
|
| 852 |
+
"errors": 0,
|
| 853 |
+
"abstained": 0,
|
| 854 |
+
"pending": 0,
|
| 855 |
+
"scored_requests": 4533,
|
| 856 |
+
"metric": "Sarcasm F1 \u00b7 track A, English",
|
| 857 |
+
"score": 0.5081,
|
| 858 |
+
"reference_same_cases": null,
|
| 859 |
+
"median_ms": 12.9,
|
| 860 |
+
"tracks": [
|
| 861 |
+
{
|
| 862 |
+
"track": "A \u00b7 Arabic",
|
| 863 |
+
"score": 0.3183,
|
| 864 |
+
"headline": false
|
| 865 |
+
},
|
| 866 |
+
{
|
| 867 |
+
"track": "A \u00b7 English",
|
| 868 |
+
"score": 0.5081,
|
| 869 |
+
"headline": true
|
| 870 |
+
},
|
| 871 |
+
{
|
| 872 |
+
"track": "C \u00b7 Arabic pairs",
|
| 873 |
+
"score": 0.7889,
|
| 874 |
+
"headline": false
|
| 875 |
+
},
|
| 876 |
+
{
|
| 877 |
+
"track": "C \u00b7 English pairs",
|
| 878 |
+
"score": 0.9497,
|
| 879 |
+
"headline": false
|
| 880 |
+
}
|
| 881 |
+
],
|
| 882 |
+
"in_index": true
|
| 883 |
+
},
|
| 884 |
+
"41": {
|
| 885 |
+
"catalog_id": 41,
|
| 886 |
+
"dataset": "VAST",
|
| 887 |
+
"requests": 2939,
|
| 888 |
+
"answered": 2939,
|
| 889 |
+
"unsupported": 0,
|
| 890 |
+
"errors": 0,
|
| 891 |
+
"abstained": 0,
|
| 892 |
+
"pending": 0,
|
| 893 |
+
"scored_requests": 2939,
|
| 894 |
+
"metric": "macro-F1",
|
| 895 |
+
"score": 0.7791,
|
| 896 |
+
"reference_same_cases": null,
|
| 897 |
+
"median_ms": 23.6,
|
| 898 |
+
"in_index": true
|
| 899 |
+
},
|
| 900 |
+
"42": {
|
| 901 |
+
"catalog_id": 42,
|
| 902 |
+
"dataset": "NLI4CT",
|
| 903 |
+
"requests": 5433,
|
| 904 |
+
"answered": 5433,
|
| 905 |
+
"unsupported": 0,
|
| 906 |
+
"errors": 0,
|
| 907 |
+
"abstained": 0,
|
| 908 |
+
"pending": 0,
|
| 909 |
+
"scored_requests": 5433,
|
| 910 |
+
"metric": "macro-F1",
|
| 911 |
+
"score": 0.8027,
|
| 912 |
+
"reference_same_cases": null,
|
| 913 |
+
"median_ms": 54.9,
|
| 914 |
+
"in_index": false
|
| 915 |
+
},
|
| 916 |
+
"43": {
|
| 917 |
+
"catalog_id": 43,
|
| 918 |
+
"dataset": "CRUXEval",
|
| 919 |
+
"requests": 503,
|
| 920 |
+
"answered": 503,
|
| 921 |
+
"unsupported": 0,
|
| 922 |
+
"errors": 0,
|
| 923 |
+
"abstained": 0,
|
| 924 |
+
"pending": 0,
|
| 925 |
+
"scored_requests": 503,
|
| 926 |
+
"metric": "accuracy",
|
| 927 |
+
"score": 0.5606,
|
| 928 |
+
"reference_same_cases": null,
|
| 929 |
+
"median_ms": 20.3,
|
| 930 |
+
"in_index": true
|
| 931 |
+
},
|
| 932 |
+
"44": {
|
| 933 |
+
"catalog_id": 44,
|
| 934 |
+
"dataset": "CLadder",
|
| 935 |
+
"requests": 4933,
|
| 936 |
+
"answered": 4933,
|
| 937 |
+
"unsupported": 0,
|
| 938 |
+
"errors": 0,
|
| 939 |
+
"abstained": 0,
|
| 940 |
+
"pending": 0,
|
| 941 |
+
"scored_requests": 4933,
|
| 942 |
+
"metric": "accuracy",
|
| 943 |
+
"score": 0.6615,
|
| 944 |
+
"reference_same_cases": null,
|
| 945 |
+
"median_ms": 19.9,
|
| 946 |
+
"in_index": true
|
| 947 |
+
},
|
| 948 |
+
"45": {
|
| 949 |
+
"catalog_id": 45,
|
| 950 |
+
"dataset": "HLE",
|
| 951 |
+
"requests": 434,
|
| 952 |
+
"answered": 434,
|
| 953 |
+
"unsupported": 0,
|
| 954 |
+
"errors": 0,
|
| 955 |
+
"abstained": 0,
|
| 956 |
+
"pending": 0,
|
| 957 |
+
"scored_requests": 434,
|
| 958 |
+
"metric": "accuracy",
|
| 959 |
+
"score": 0.1106,
|
| 960 |
+
"reference_same_cases": null,
|
| 961 |
+
"median_ms": 30.4,
|
| 962 |
+
"in_index": false
|
| 963 |
+
},
|
| 964 |
+
"48": {
|
| 965 |
+
"catalog_id": 48,
|
| 966 |
+
"dataset": "ForecastBench",
|
| 967 |
+
"requests": 10072,
|
| 968 |
+
"answered": 10072,
|
| 969 |
+
"unsupported": 0,
|
| 970 |
+
"errors": 0,
|
| 971 |
+
"abstained": 0,
|
| 972 |
+
"pending": 0,
|
| 973 |
+
"scored_requests": 10072,
|
| 974 |
+
"metric": "Brier (lower is better)",
|
| 975 |
+
"score": 0.1792,
|
| 976 |
+
"reference_same_cases": null,
|
| 977 |
+
"median_ms": 73.2,
|
| 978 |
+
"in_index": false
|
| 979 |
+
},
|
| 980 |
+
"50": {
|
| 981 |
+
"catalog_id": 50,
|
| 982 |
+
"dataset": "Habermas Machine",
|
| 983 |
+
"requests": 1609,
|
| 984 |
+
"answered": 1609,
|
| 985 |
+
"unsupported": 0,
|
| 986 |
+
"errors": 0,
|
| 987 |
+
"abstained": 0,
|
| 988 |
+
"pending": 0,
|
| 989 |
+
"scored_requests": 1609,
|
| 990 |
+
"metric": "accuracy",
|
| 991 |
+
"score": 0.4599,
|
| 992 |
+
"reference_same_cases": null,
|
| 993 |
+
"median_ms": 81.2,
|
| 994 |
+
"in_index": true
|
| 995 |
+
}
|
| 996 |
+
},
|
| 997 |
+
"frozen_panel": {
|
| 998 |
+
"scores": {
|
| 999 |
+
"balanced_raw": 43.89,
|
| 1000 |
+
"balanced_skill": 31.69,
|
| 1001 |
+
"breadth_skill": 27.41
|
| 1002 |
+
},
|
| 1003 |
+
"categories": [
|
| 1004 |
+
{
|
| 1005 |
+
"id": "knowledge",
|
| 1006 |
+
"raw": 0.615,
|
| 1007 |
+
"skill": 0.4278,
|
| 1008 |
+
"coverage": 1.0,
|
| 1009 |
+
"pending": 0.0
|
| 1010 |
+
},
|
| 1011 |
+
{
|
| 1012 |
+
"id": "language",
|
| 1013 |
+
"raw": 0.5883,
|
| 1014 |
+
"skill": 0.4886,
|
| 1015 |
+
"coverage": 1.0,
|
| 1016 |
+
"pending": 0.0
|
| 1017 |
+
},
|
| 1018 |
+
{
|
| 1019 |
+
"id": "tools",
|
| 1020 |
+
"raw": 0.4227,
|
| 1021 |
+
"skill": 0.3486,
|
| 1022 |
+
"coverage": 0.6,
|
| 1023 |
+
"pending": 0.4
|
| 1024 |
+
},
|
| 1025 |
+
{
|
| 1026 |
+
"id": "games",
|
| 1027 |
+
"raw": 0.0458,
|
| 1028 |
+
"skill": 0.032,
|
| 1029 |
+
"coverage": 0.2,
|
| 1030 |
+
"pending": 0.8
|
| 1031 |
+
},
|
| 1032 |
+
{
|
| 1033 |
+
"id": "arts",
|
| 1034 |
+
"raw": 0.5225,
|
| 1035 |
+
"skill": 0.2876,
|
| 1036 |
+
"coverage": 1.0,
|
| 1037 |
+
"pending": 0.0
|
| 1038 |
+
}
|
| 1039 |
+
],
|
| 1040 |
+
"coverage": 0.76,
|
| 1041 |
+
"pending": 0.24,
|
| 1042 |
+
"note": "25-benchmark frozen panel with the six interactive benchmarks scored as zero (provisional lower bound); not the headline index."
|
| 1043 |
+
},
|
| 1044 |
+
"formulas": [
|
| 1045 |
+
{
|
| 1046 |
+
"id": "balanced_raw",
|
| 1047 |
+
"normalization": "raw",
|
| 1048 |
+
"formula": "100*sum_c(0.2*category[c])"
|
| 1049 |
+
},
|
| 1050 |
+
{
|
| 1051 |
+
"id": "balanced_skill",
|
| 1052 |
+
"normalization": "skill",
|
| 1053 |
+
"formula": "100*sum_c(0.2*category[c])"
|
| 1054 |
+
},
|
| 1055 |
+
{
|
| 1056 |
+
"id": "breadth_skill",
|
| 1057 |
+
"normalization": "skill",
|
| 1058 |
+
"shift": 0.1,
|
| 1059 |
+
"formula": "100*(product_c((0.1+0.9*category[c])**0.2)-0.1)/0.9"
|
| 1060 |
+
}
|
| 1061 |
+
],
|
| 1062 |
+
"panel_id": "core25-observed-protocol-v1",
|
| 1063 |
+
"edition": "Decision Index 0.1 (archived 19-benchmark edition)",
|
| 1064 |
+
"subset": {
|
| 1065 |
+
"name": "0.1 suite minus the 3,000-request DI-S selection sample",
|
| 1066 |
+
"requests": 129422,
|
| 1067 |
+
"ids_sha256": "b738768b0cd2076ead8c637cddd2d1b233094f46524153da16444ff7a98f6aab",
|
| 1068 |
+
"note": "The 'suite' block above describes the parent 0.1 suite; scores cover the subset only."
|
| 1069 |
+
}
|
| 1070 |
+
}
|
eval/jevbench-public-proxy.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"tiers": {
|
| 3 |
+
"easy": 1.0,
|
| 4 |
+
"standard": 0.9722222222222222,
|
| 5 |
+
"hard": 0.6936936936936937
|
| 6 |
+
},
|
| 7 |
+
"hard_ece": 0.13599638226455113,
|
| 8 |
+
"prob_tvd": 0.205386479985251,
|
| 9 |
+
"calibration": 76.13103777428233,
|
| 10 |
+
"intelligence_public": 79.18539489648231,
|
| 11 |
+
"input_tokens_per_decision": 694.922077922078,
|
| 12 |
+
"temperature": 1.0,
|
| 13 |
+
"hard_by_family": {
|
| 14 |
+
"adversarial": 1.0,
|
| 15 |
+
"ambiguous": 1.0,
|
| 16 |
+
"judge_hard": 0.647,
|
| 17 |
+
"long_policy": 0.579,
|
| 18 |
+
"multi_hop": 0.722,
|
| 19 |
+
"probability": 0.8,
|
| 20 |
+
"routing_hard": 1.0,
|
| 21 |
+
"temporal_numeric": 0.333,
|
| 22 |
+
"tradeoff": 0.5,
|
| 23 |
+
"trap": 1.0
|
| 24 |
+
},
|
| 25 |
+
"p50_s": 0.06154986098408699,
|
| 26 |
+
"p95_s": 0.2521260780049488,
|
| 27 |
+
"speed": 87.47933579517849,
|
| 28 |
+
"note": "JevBench v1.2 public items (231) scored with JevBench's own code; a development proxy, not an official JevBench result."
|
| 29 |
+
}
|
generation_config.json
ADDED
|
@@ -0,0 +1,10 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"eos_token_id": [
|
| 3 |
+
248046,
|
| 4 |
+
248044
|
| 5 |
+
],
|
| 6 |
+
"do_sample": true,
|
| 7 |
+
"temperature": 0.6,
|
| 8 |
+
"top_k": 20,
|
| 9 |
+
"top_p": 0.95
|
| 10 |
+
}
|
merges.txt
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
model-00001-of-00004.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9da483e4c921f161562d994709fe7181a05d1a14692026d9943da80021247d82
|
| 3 |
+
size 4624015664
|
model-00002-of-00004.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:efbf9af22c3f00f32289c50f9cd9ed5abc6cf9ab4af957d9adae1c80cd0daeae
|
| 3 |
+
size 3911711088
|
model-00003-of-00004.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:3fb71b091963d851c95352c3d3a24ecefc114a3b281e916943a129f173db05d7
|
| 3 |
+
size 5360262952
|
model-00004-of-00004.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6c8a6365689c40acdf1c7605819323cd09f7f5f8818497e9cd56f50b7ef38d34
|
| 3 |
+
size 4923731464
|
model.safetensors.index.json
ADDED
|
@@ -0,0 +1,767 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"metadata": {
|
| 3 |
+
"total_size": 18819627488
|
| 4 |
+
},
|
| 5 |
+
"weight_map": {
|
| 6 |
+
"model.visual.blocks.8.norm2.weight": "model-00001-of-00004.safetensors",
|
| 7 |
+
"model.language_model.layers.26.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
|
| 8 |
+
"model.language_model.layers.26.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
|
| 9 |
+
"model.visual.blocks.2.attn.proj.bias": "model-00001-of-00004.safetensors",
|
| 10 |
+
"model.language_model.layers.0.linear_attn.dt_bias": "model-00001-of-00004.safetensors",
|
| 11 |
+
"model.language_model.layers.10.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
|
| 12 |
+
"model.language_model.layers.1.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 13 |
+
"model.visual.blocks.11.norm2.weight": "model-00001-of-00004.safetensors",
|
| 14 |
+
"model.visual.blocks.14.norm2.weight": "model-00001-of-00004.safetensors",
|
| 15 |
+
"model.language_model.layers.27.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
|
| 16 |
+
"model.language_model.layers.13.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 17 |
+
"model.language_model.layers.22.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 18 |
+
"model.visual.blocks.9.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 19 |
+
"model.visual.blocks.5.norm2.bias": "model-00001-of-00004.safetensors",
|
| 20 |
+
"model.visual.blocks.22.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 21 |
+
"model.language_model.layers.31.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
|
| 22 |
+
"model.visual.blocks.2.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 23 |
+
"model.visual.blocks.14.mlp.linear_fc1.bias": "model-00001-of-00004.safetensors",
|
| 24 |
+
"model.language_model.layers.14.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
|
| 25 |
+
"model.visual.blocks.3.mlp.linear_fc1.weight": "model-00001-of-00004.safetensors",
|
| 26 |
+
"model.language_model.layers.4.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
|
| 27 |
+
"model.visual.blocks.14.attn.proj.weight": "model-00001-of-00004.safetensors",
|
| 28 |
+
"model.language_model.layers.20.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 29 |
+
"model.language_model.layers.1.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 30 |
+
"model.visual.merger.linear_fc2.weight": "model-00001-of-00004.safetensors",
|
| 31 |
+
"model.visual.blocks.2.mlp.linear_fc2.weight": "model-00001-of-00004.safetensors",
|
| 32 |
+
"model.language_model.layers.13.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
|
| 33 |
+
"model.language_model.layers.1.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 34 |
+
"model.visual.blocks.9.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 35 |
+
"model.language_model.layers.6.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 36 |
+
"model.language_model.layers.29.linear_attn.A_log": "model-00001-of-00004.safetensors",
|
| 37 |
+
"model.language_model.layers.9.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
|
| 38 |
+
"model.language_model.layers.26.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
|
| 39 |
+
"model.language_model.layers.6.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 40 |
+
"model.visual.blocks.4.norm1.bias": "model-00001-of-00004.safetensors",
|
| 41 |
+
"model.visual.blocks.12.attn.proj.weight": "model-00001-of-00004.safetensors",
|
| 42 |
+
"model.visual.blocks.20.norm1.bias": "model-00001-of-00004.safetensors",
|
| 43 |
+
"model.visual.blocks.5.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 44 |
+
"model.language_model.layers.7.self_attn.q_norm.weight": "model-00001-of-00004.safetensors",
|
| 45 |
+
"model.language_model.layers.19.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
|
| 46 |
+
"model.visual.blocks.17.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 47 |
+
"model.language_model.layers.19.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 48 |
+
"model.visual.blocks.26.norm2.bias": "model-00001-of-00004.safetensors",
|
| 49 |
+
"model.language_model.layers.24.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
|
| 50 |
+
"model.language_model.layers.5.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
|
| 51 |
+
"model.language_model.layers.8.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
|
| 52 |
+
"model.language_model.layers.12.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
|
| 53 |
+
"model.visual.blocks.21.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 54 |
+
"model.visual.blocks.13.mlp.linear_fc1.bias": "model-00001-of-00004.safetensors",
|
| 55 |
+
"model.language_model.layers.9.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
|
| 56 |
+
"model.visual.merger.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 57 |
+
"model.visual.blocks.11.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 58 |
+
"model.language_model.layers.18.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 59 |
+
"model.visual.blocks.20.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 60 |
+
"model.visual.blocks.12.norm2.bias": "model-00001-of-00004.safetensors",
|
| 61 |
+
"model.visual.blocks.15.norm2.bias": "model-00001-of-00004.safetensors",
|
| 62 |
+
"model.visual.blocks.0.mlp.linear_fc1.bias": "model-00001-of-00004.safetensors",
|
| 63 |
+
"model.visual.blocks.10.norm1.weight": "model-00001-of-00004.safetensors",
|
| 64 |
+
"model.language_model.layers.18.linear_attn.A_log": "model-00001-of-00004.safetensors",
|
| 65 |
+
"model.language_model.layers.11.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 66 |
+
"model.language_model.layers.15.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 67 |
+
"model.visual.blocks.15.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 68 |
+
"model.language_model.layers.17.linear_attn.A_log": "model-00001-of-00004.safetensors",
|
| 69 |
+
"model.language_model.layers.9.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 70 |
+
"model.visual.blocks.19.attn.proj.bias": "model-00001-of-00004.safetensors",
|
| 71 |
+
"model.language_model.layers.3.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 72 |
+
"model.visual.blocks.17.attn.proj.bias": "model-00001-of-00004.safetensors",
|
| 73 |
+
"model.language_model.layers.9.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
|
| 74 |
+
"model.visual.blocks.18.norm1.weight": "model-00001-of-00004.safetensors",
|
| 75 |
+
"model.visual.blocks.8.norm1.bias": "model-00001-of-00004.safetensors",
|
| 76 |
+
"model.language_model.layers.18.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
|
| 77 |
+
"model.visual.blocks.3.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 78 |
+
"model.language_model.layers.31.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
|
| 79 |
+
"model.language_model.layers.9.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
|
| 80 |
+
"model.visual.blocks.12.attn.qkv.bias": "model-00001-of-00004.safetensors",
|
| 81 |
+
"model.visual.blocks.0.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 82 |
+
"model.visual.pos_embed.weight": "model-00001-of-00004.safetensors",
|
| 83 |
+
"model.language_model.layers.7.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
|
| 84 |
+
"model.language_model.layers.11.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
|
| 85 |
+
"model.visual.blocks.22.norm2.bias": "model-00001-of-00004.safetensors",
|
| 86 |
+
"model.visual.blocks.21.norm2.weight": "model-00001-of-00004.safetensors",
|
| 87 |
+
"model.language_model.layers.18.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
|
| 88 |
+
"model.visual.blocks.20.mlp.linear_fc1.bias": "model-00001-of-00004.safetensors",
|
| 89 |
+
"model.language_model.layers.3.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
|
| 90 |
+
"model.language_model.layers.29.linear_attn.dt_bias": "model-00001-of-00004.safetensors",
|
| 91 |
+
"model.language_model.layers.29.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
|
| 92 |
+
"model.language_model.layers.15.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
|
| 93 |
+
"model.language_model.layers.0.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 94 |
+
"model.language_model.layers.23.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
|
| 95 |
+
"model.language_model.layers.19.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
|
| 96 |
+
"model.visual.blocks.9.norm1.weight": "model-00001-of-00004.safetensors",
|
| 97 |
+
"model.language_model.layers.30.linear_attn.A_log": "model-00001-of-00004.safetensors",
|
| 98 |
+
"model.visual.blocks.0.norm2.weight": "model-00001-of-00004.safetensors",
|
| 99 |
+
"model.language_model.layers.6.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 100 |
+
"model.visual.blocks.0.norm2.bias": "model-00001-of-00004.safetensors",
|
| 101 |
+
"model.visual.patch_embed.proj.bias": "model-00001-of-00004.safetensors",
|
| 102 |
+
"model.language_model.layers.13.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
|
| 103 |
+
"model.language_model.layers.12.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 104 |
+
"model.visual.blocks.4.mlp.linear_fc1.weight": "model-00001-of-00004.safetensors",
|
| 105 |
+
"model.language_model.layers.7.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
|
| 106 |
+
"model.language_model.layers.23.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 107 |
+
"model.visual.blocks.18.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 108 |
+
"model.language_model.layers.10.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 109 |
+
"model.visual.blocks.10.norm2.weight": "model-00001-of-00004.safetensors",
|
| 110 |
+
"model.language_model.layers.6.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
|
| 111 |
+
"model.language_model.layers.4.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 112 |
+
"model.language_model.layers.16.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
|
| 113 |
+
"model.language_model.layers.28.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 114 |
+
"model.language_model.layers.15.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
|
| 115 |
+
"model.visual.blocks.15.norm2.weight": "model-00001-of-00004.safetensors",
|
| 116 |
+
"model.visual.blocks.26.attn.qkv.bias": "model-00001-of-00004.safetensors",
|
| 117 |
+
"model.language_model.layers.7.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
|
| 118 |
+
"model.visual.blocks.9.norm2.weight": "model-00001-of-00004.safetensors",
|
| 119 |
+
"model.language_model.layers.21.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
|
| 120 |
+
"model.language_model.layers.22.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
|
| 121 |
+
"model.visual.blocks.5.mlp.linear_fc1.weight": "model-00001-of-00004.safetensors",
|
| 122 |
+
"model.language_model.layers.22.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 123 |
+
"model.language_model.layers.6.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
|
| 124 |
+
"model.language_model.layers.24.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 125 |
+
"model.visual.blocks.25.mlp.linear_fc1.bias": "model-00001-of-00004.safetensors",
|
| 126 |
+
"model.language_model.layers.18.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
|
| 127 |
+
"model.visual.blocks.20.mlp.linear_fc1.weight": "model-00001-of-00004.safetensors",
|
| 128 |
+
"model.visual.blocks.23.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 129 |
+
"model.visual.blocks.8.attn.proj.weight": "model-00001-of-00004.safetensors",
|
| 130 |
+
"model.language_model.layers.26.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 131 |
+
"model.language_model.layers.13.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
|
| 132 |
+
"model.language_model.layers.20.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 133 |
+
"model.visual.blocks.10.norm1.bias": "model-00001-of-00004.safetensors",
|
| 134 |
+
"model.language_model.layers.15.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 135 |
+
"model.language_model.layers.31.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 136 |
+
"model.visual.blocks.22.mlp.linear_fc1.bias": "model-00001-of-00004.safetensors",
|
| 137 |
+
"model.language_model.layers.14.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
|
| 138 |
+
"model.visual.blocks.25.norm2.bias": "model-00001-of-00004.safetensors",
|
| 139 |
+
"model.visual.blocks.14.attn.proj.bias": "model-00001-of-00004.safetensors",
|
| 140 |
+
"model.language_model.layers.30.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
|
| 141 |
+
"model.language_model.layers.21.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
|
| 142 |
+
"model.language_model.layers.4.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
|
| 143 |
+
"model.language_model.layers.12.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
|
| 144 |
+
"model.language_model.layers.1.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
|
| 145 |
+
"model.language_model.layers.21.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 146 |
+
"model.language_model.layers.24.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
|
| 147 |
+
"model.visual.blocks.23.attn.proj.weight": "model-00001-of-00004.safetensors",
|
| 148 |
+
"model.language_model.layers.20.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 149 |
+
"model.visual.merger.norm.weight": "model-00001-of-00004.safetensors",
|
| 150 |
+
"model.language_model.layers.13.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
|
| 151 |
+
"model.visual.blocks.20.attn.proj.bias": "model-00001-of-00004.safetensors",
|
| 152 |
+
"model.language_model.layers.11.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
|
| 153 |
+
"model.language_model.layers.18.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 154 |
+
"model.language_model.layers.21.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 155 |
+
"model.language_model.layers.30.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
|
| 156 |
+
"model.visual.blocks.1.norm1.weight": "model-00001-of-00004.safetensors",
|
| 157 |
+
"model.visual.blocks.14.norm1.weight": "model-00001-of-00004.safetensors",
|
| 158 |
+
"model.language_model.layers.13.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
|
| 159 |
+
"model.visual.blocks.10.mlp.linear_fc1.weight": "model-00001-of-00004.safetensors",
|
| 160 |
+
"model.language_model.layers.24.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 161 |
+
"model.language_model.layers.27.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 162 |
+
"model.visual.blocks.3.norm1.weight": "model-00001-of-00004.safetensors",
|
| 163 |
+
"model.language_model.layers.1.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 164 |
+
"model.language_model.layers.23.self_attn.q_norm.weight": "model-00001-of-00004.safetensors",
|
| 165 |
+
"model.visual.merger.linear_fc1.weight": "model-00001-of-00004.safetensors",
|
| 166 |
+
"model.language_model.layers.8.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 167 |
+
"model.language_model.layers.1.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
|
| 168 |
+
"model.language_model.layers.25.linear_attn.A_log": "model-00001-of-00004.safetensors",
|
| 169 |
+
"model.language_model.layers.17.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
|
| 170 |
+
"model.language_model.layers.13.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
|
| 171 |
+
"model.visual.blocks.1.norm2.bias": "model-00001-of-00004.safetensors",
|
| 172 |
+
"model.visual.blocks.16.attn.proj.bias": "model-00001-of-00004.safetensors",
|
| 173 |
+
"model.visual.blocks.24.norm1.weight": "model-00001-of-00004.safetensors",
|
| 174 |
+
"model.language_model.layers.3.self_attn.q_norm.weight": "model-00001-of-00004.safetensors",
|
| 175 |
+
"model.visual.blocks.5.norm1.bias": "model-00001-of-00004.safetensors",
|
| 176 |
+
"model.language_model.layers.0.linear_attn.A_log": "model-00001-of-00004.safetensors",
|
| 177 |
+
"model.language_model.layers.28.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 178 |
+
"model.language_model.layers.14.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 179 |
+
"model.visual.blocks.3.attn.proj.bias": "model-00001-of-00004.safetensors",
|
| 180 |
+
"model.language_model.layers.16.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 181 |
+
"model.language_model.layers.24.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
|
| 182 |
+
"model.visual.blocks.20.attn.proj.weight": "model-00001-of-00004.safetensors",
|
| 183 |
+
"model.visual.blocks.4.norm1.weight": "model-00001-of-00004.safetensors",
|
| 184 |
+
"model.language_model.layers.19.self_attn.q_norm.weight": "model-00001-of-00004.safetensors",
|
| 185 |
+
"model.visual.blocks.8.mlp.linear_fc2.weight": "model-00001-of-00004.safetensors",
|
| 186 |
+
"model.language_model.layers.2.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 187 |
+
"model.visual.blocks.24.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 188 |
+
"model.visual.blocks.23.attn.proj.bias": "model-00001-of-00004.safetensors",
|
| 189 |
+
"model.language_model.layers.30.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 190 |
+
"model.language_model.layers.18.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 191 |
+
"model.visual.blocks.6.attn.proj.weight": "model-00001-of-00004.safetensors",
|
| 192 |
+
"model.language_model.layers.6.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 193 |
+
"model.visual.blocks.8.norm2.bias": "model-00001-of-00004.safetensors",
|
| 194 |
+
"model.visual.blocks.3.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 195 |
+
"model.language_model.layers.10.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 196 |
+
"model.visual.blocks.17.mlp.linear_fc2.weight": "model-00001-of-00004.safetensors",
|
| 197 |
+
"model.language_model.layers.22.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
|
| 198 |
+
"model.language_model.layers.21.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 199 |
+
"model.language_model.layers.10.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
|
| 200 |
+
"model.language_model.layers.9.linear_attn.dt_bias": "model-00001-of-00004.safetensors",
|
| 201 |
+
"model.language_model.layers.4.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 202 |
+
"model.language_model.layers.26.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
|
| 203 |
+
"model.visual.blocks.21.norm1.weight": "model-00001-of-00004.safetensors",
|
| 204 |
+
"model.visual.blocks.13.norm1.weight": "model-00001-of-00004.safetensors",
|
| 205 |
+
"model.language_model.layers.0.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
|
| 206 |
+
"model.visual.blocks.7.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 207 |
+
"model.language_model.layers.8.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 208 |
+
"model.language_model.layers.31.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
|
| 209 |
+
"model.visual.blocks.3.mlp.linear_fc1.bias": "model-00001-of-00004.safetensors",
|
| 210 |
+
"model.visual.blocks.15.attn.qkv.bias": "model-00001-of-00004.safetensors",
|
| 211 |
+
"model.language_model.layers.8.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
|
| 212 |
+
"model.language_model.layers.0.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 213 |
+
"model.language_model.layers.13.linear_attn.A_log": "model-00001-of-00004.safetensors",
|
| 214 |
+
"model.visual.blocks.26.mlp.linear_fc1.weight": "model-00001-of-00004.safetensors",
|
| 215 |
+
"model.language_model.layers.6.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
|
| 216 |
+
"model.visual.blocks.14.mlp.linear_fc2.weight": "model-00001-of-00004.safetensors",
|
| 217 |
+
"model.visual.blocks.1.norm1.bias": "model-00001-of-00004.safetensors",
|
| 218 |
+
"model.visual.blocks.7.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 219 |
+
"model.language_model.layers.1.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
|
| 220 |
+
"model.language_model.layers.4.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 221 |
+
"model.visual.blocks.1.norm2.weight": "model-00001-of-00004.safetensors",
|
| 222 |
+
"model.language_model.layers.0.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
|
| 223 |
+
"model.language_model.layers.8.linear_attn.dt_bias": "model-00001-of-00004.safetensors",
|
| 224 |
+
"model.visual.blocks.25.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 225 |
+
"model.language_model.layers.25.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
|
| 226 |
+
"model.visual.blocks.23.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 227 |
+
"model.language_model.layers.4.linear_attn.A_log": "model-00001-of-00004.safetensors",
|
| 228 |
+
"model.visual.blocks.25.mlp.linear_fc1.weight": "model-00001-of-00004.safetensors",
|
| 229 |
+
"model.visual.blocks.8.norm1.weight": "model-00001-of-00004.safetensors",
|
| 230 |
+
"model.visual.blocks.1.mlp.linear_fc2.weight": "model-00001-of-00004.safetensors",
|
| 231 |
+
"model.visual.blocks.8.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 232 |
+
"model.language_model.layers.22.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 233 |
+
"model.language_model.layers.24.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
|
| 234 |
+
"model.language_model.layers.30.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
|
| 235 |
+
"model.language_model.layers.22.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
|
| 236 |
+
"model.language_model.layers.3.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
|
| 237 |
+
"model.visual.blocks.5.norm1.weight": "model-00001-of-00004.safetensors",
|
| 238 |
+
"model.language_model.layers.0.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 239 |
+
"model.visual.blocks.14.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 240 |
+
"model.language_model.layers.18.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
|
| 241 |
+
"model.language_model.layers.8.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 242 |
+
"model.visual.blocks.22.attn.qkv.bias": "model-00001-of-00004.safetensors",
|
| 243 |
+
"model.language_model.layers.14.linear_attn.dt_bias": "model-00001-of-00004.safetensors",
|
| 244 |
+
"model.language_model.layers.17.input_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 245 |
+
"model.language_model.layers.15.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
|
| 246 |
+
"model.visual.blocks.13.attn.proj.bias": "model-00001-of-00004.safetensors",
|
| 247 |
+
"model.visual.blocks.14.norm2.bias": "model-00001-of-00004.safetensors",
|
| 248 |
+
"model.visual.blocks.9.attn.proj.weight": "model-00001-of-00004.safetensors",
|
| 249 |
+
"model.language_model.layers.5.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
|
| 250 |
+
"model.visual.blocks.25.attn.qkv.weight": "model-00001-of-00004.safetensors",
|
| 251 |
+
"model.visual.blocks.7.norm1.bias": "model-00001-of-00004.safetensors",
|
| 252 |
+
"model.language_model.layers.5.linear_attn.A_log": "model-00001-of-00004.safetensors",
|
| 253 |
+
"model.visual.blocks.17.norm2.bias": "model-00001-of-00004.safetensors",
|
| 254 |
+
"model.visual.blocks.25.norm1.bias": "model-00001-of-00004.safetensors",
|
| 255 |
+
"model.language_model.layers.28.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
|
| 256 |
+
"model.visual.blocks.5.mlp.linear_fc2.weight": "model-00001-of-00004.safetensors",
|
| 257 |
+
"model.language_model.layers.2.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
|
| 258 |
+
"model.language_model.layers.17.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
|
| 259 |
+
"model.language_model.layers.8.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
|
| 260 |
+
"model.visual.blocks.22.mlp.linear_fc1.weight": "model-00001-of-00004.safetensors",
|
| 261 |
+
"model.visual.blocks.3.norm2.bias": "model-00001-of-00004.safetensors",
|
| 262 |
+
"model.visual.blocks.2.norm2.bias": "model-00001-of-00004.safetensors",
|
| 263 |
+
"model.language_model.layers.12.linear_attn.dt_bias": "model-00001-of-00004.safetensors",
|
| 264 |
+
"model.language_model.layers.31.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
|
| 265 |
+
"model.language_model.layers.27.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
|
| 266 |
+
"model.language_model.layers.16.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
|
| 267 |
+
"model.visual.blocks.24.attn.proj.bias": "model-00001-of-00004.safetensors",
|
| 268 |
+
"model.visual.blocks.10.attn.qkv.bias": "model-00001-of-00004.safetensors",
|
| 269 |
+
"model.visual.blocks.12.mlp.linear_fc2.bias": "model-00001-of-00004.safetensors",
|
| 270 |
+
"model.language_model.layers.21.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 271 |
+
"model.language_model.layers.13.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
|
| 272 |
+
"model.language_model.embed_tokens.weight": "model-00002-of-00004.safetensors",
|
| 273 |
+
"model.language_model.layers.2.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
|
| 274 |
+
"model.language_model.layers.28.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
|
| 275 |
+
"model.visual.blocks.7.mlp.linear_fc1.bias": "model-00002-of-00004.safetensors",
|
| 276 |
+
"model.visual.blocks.26.norm1.weight": "model-00002-of-00004.safetensors",
|
| 277 |
+
"model.visual.blocks.7.attn.qkv.bias": "model-00002-of-00004.safetensors",
|
| 278 |
+
"model.language_model.layers.28.linear_attn.A_log": "model-00002-of-00004.safetensors",
|
| 279 |
+
"model.visual.blocks.19.attn.proj.weight": "model-00002-of-00004.safetensors",
|
| 280 |
+
"model.visual.blocks.15.attn.proj.weight": "model-00002-of-00004.safetensors",
|
| 281 |
+
"model.language_model.layers.16.linear_attn.in_proj_z.weight": "model-00002-of-00004.safetensors",
|
| 282 |
+
"model.visual.blocks.17.attn.proj.weight": "model-00002-of-00004.safetensors",
|
| 283 |
+
"model.visual.blocks.14.attn.qkv.weight": "model-00002-of-00004.safetensors",
|
| 284 |
+
"model.visual.blocks.15.norm1.bias": "model-00002-of-00004.safetensors",
|
| 285 |
+
"model.visual.blocks.7.attn.proj.bias": "model-00002-of-00004.safetensors",
|
| 286 |
+
"model.language_model.layers.3.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
|
| 287 |
+
"model.language_model.layers.20.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
|
| 288 |
+
"model.language_model.layers.17.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
|
| 289 |
+
"model.language_model.layers.23.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
|
| 290 |
+
"model.language_model.layers.16.linear_attn.A_log": "model-00002-of-00004.safetensors",
|
| 291 |
+
"model.language_model.layers.0.linear_attn.conv1d.weight": "model-00002-of-00004.safetensors",
|
| 292 |
+
"model.visual.blocks.21.attn.qkv.weight": "model-00002-of-00004.safetensors",
|
| 293 |
+
"model.visual.blocks.6.mlp.linear_fc2.bias": "model-00002-of-00004.safetensors",
|
| 294 |
+
"model.visual.blocks.25.norm2.weight": "model-00002-of-00004.safetensors",
|
| 295 |
+
"model.language_model.layers.2.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
|
| 296 |
+
"model.language_model.layers.2.linear_attn.in_proj_a.weight": "model-00002-of-00004.safetensors",
|
| 297 |
+
"model.language_model.layers.10.linear_attn.in_proj_b.weight": "model-00002-of-00004.safetensors",
|
| 298 |
+
"model.language_model.layers.25.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
|
| 299 |
+
"model.visual.blocks.12.norm1.bias": "model-00002-of-00004.safetensors",
|
| 300 |
+
"model.visual.blocks.6.attn.qkv.bias": "model-00002-of-00004.safetensors",
|
| 301 |
+
"model.visual.blocks.24.norm1.bias": "model-00002-of-00004.safetensors",
|
| 302 |
+
"model.language_model.layers.26.linear_attn.A_log": "model-00002-of-00004.safetensors",
|
| 303 |
+
"model.language_model.layers.6.linear_attn.A_log": "model-00002-of-00004.safetensors",
|
| 304 |
+
"model.visual.patch_embed.proj.weight": "model-00002-of-00004.safetensors",
|
| 305 |
+
"model.language_model.layers.19.input_layernorm.weight": "model-00002-of-00004.safetensors",
|
| 306 |
+
"model.language_model.layers.9.linear_attn.in_proj_a.weight": "model-00002-of-00004.safetensors",
|
| 307 |
+
"model.language_model.layers.11.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
|
| 308 |
+
"model.language_model.layers.30.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
|
| 309 |
+
"model.language_model.layers.10.linear_attn.A_log": "model-00002-of-00004.safetensors",
|
| 310 |
+
"model.language_model.layers.29.linear_attn.in_proj_a.weight": "model-00002-of-00004.safetensors",
|
| 311 |
+
"model.language_model.layers.12.linear_attn.conv1d.weight": "model-00002-of-00004.safetensors",
|
| 312 |
+
"model.visual.blocks.4.norm2.weight": "model-00002-of-00004.safetensors",
|
| 313 |
+
"model.language_model.layers.28.linear_attn.norm.weight": "model-00002-of-00004.safetensors",
|
| 314 |
+
"model.visual.blocks.23.norm2.weight": "model-00002-of-00004.safetensors",
|
| 315 |
+
"model.visual.blocks.17.attn.qkv.bias": "model-00002-of-00004.safetensors",
|
| 316 |
+
"model.visual.blocks.17.norm1.weight": "model-00002-of-00004.safetensors",
|
| 317 |
+
"model.language_model.layers.28.linear_attn.in_proj_b.weight": "model-00002-of-00004.safetensors",
|
| 318 |
+
"model.visual.blocks.13.mlp.linear_fc2.weight": "model-00002-of-00004.safetensors",
|
| 319 |
+
"model.visual.blocks.18.mlp.linear_fc1.weight": "model-00002-of-00004.safetensors",
|
| 320 |
+
"model.language_model.layers.26.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
|
| 321 |
+
"model.visual.blocks.2.norm1.bias": "model-00002-of-00004.safetensors",
|
| 322 |
+
"model.visual.blocks.13.norm1.bias": "model-00002-of-00004.safetensors",
|
| 323 |
+
"model.language_model.layers.10.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
|
| 324 |
+
"model.language_model.layers.5.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
|
| 325 |
+
"model.language_model.layers.9.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
|
| 326 |
+
"model.visual.blocks.16.mlp.linear_fc2.weight": "model-00002-of-00004.safetensors",
|
| 327 |
+
"model.visual.blocks.16.mlp.linear_fc2.bias": "model-00002-of-00004.safetensors",
|
| 328 |
+
"model.language_model.layers.12.linear_attn.norm.weight": "model-00002-of-00004.safetensors",
|
| 329 |
+
"model.language_model.layers.25.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
|
| 330 |
+
"model.visual.blocks.12.attn.proj.bias": "model-00002-of-00004.safetensors",
|
| 331 |
+
"model.visual.blocks.6.norm1.weight": "model-00002-of-00004.safetensors",
|
| 332 |
+
"model.visual.blocks.22.mlp.linear_fc2.weight": "model-00002-of-00004.safetensors",
|
| 333 |
+
"model.visual.blocks.21.attn.proj.bias": "model-00002-of-00004.safetensors",
|
| 334 |
+
"model.language_model.layers.20.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
|
| 335 |
+
"model.visual.blocks.16.attn.qkv.weight": "model-00002-of-00004.safetensors",
|
| 336 |
+
"model.visual.blocks.22.attn.qkv.weight": "model-00002-of-00004.safetensors",
|
| 337 |
+
"model.visual.blocks.4.norm2.bias": "model-00002-of-00004.safetensors",
|
| 338 |
+
"model.language_model.layers.25.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
|
| 339 |
+
"model.language_model.layers.14.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
|
| 340 |
+
"model.language_model.layers.3.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
|
| 341 |
+
"model.language_model.layers.10.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
|
| 342 |
+
"model.visual.blocks.3.attn.proj.weight": "model-00002-of-00004.safetensors",
|
| 343 |
+
"model.language_model.layers.25.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
|
| 344 |
+
"model.language_model.layers.24.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
|
| 345 |
+
"model.visual.blocks.18.attn.qkv.bias": "model-00002-of-00004.safetensors",
|
| 346 |
+
"model.language_model.layers.21.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
|
| 347 |
+
"model.visual.blocks.16.mlp.linear_fc1.bias": "model-00002-of-00004.safetensors",
|
| 348 |
+
"model.visual.blocks.6.mlp.linear_fc2.weight": "model-00002-of-00004.safetensors",
|
| 349 |
+
"model.language_model.layers.25.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
|
| 350 |
+
"model.language_model.layers.11.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
|
| 351 |
+
"model.visual.blocks.9.attn.proj.bias": "model-00002-of-00004.safetensors",
|
| 352 |
+
"model.visual.blocks.10.mlp.linear_fc2.weight": "model-00002-of-00004.safetensors",
|
| 353 |
+
"model.language_model.layers.11.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
|
| 354 |
+
"model.visual.blocks.10.mlp.linear_fc2.bias": "model-00002-of-00004.safetensors",
|
| 355 |
+
"model.language_model.layers.21.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
|
| 356 |
+
"model.visual.blocks.9.mlp.linear_fc1.weight": "model-00002-of-00004.safetensors",
|
| 357 |
+
"model.visual.blocks.4.mlp.linear_fc2.weight": "model-00002-of-00004.safetensors",
|
| 358 |
+
"model.visual.blocks.2.attn.qkv.bias": "model-00002-of-00004.safetensors",
|
| 359 |
+
"model.language_model.layers.19.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
|
| 360 |
+
"model.visual.blocks.19.norm2.weight": "model-00002-of-00004.safetensors",
|
| 361 |
+
"model.visual.blocks.11.norm1.weight": "model-00002-of-00004.safetensors",
|
| 362 |
+
"model.visual.blocks.12.mlp.linear_fc1.weight": "model-00002-of-00004.safetensors",
|
| 363 |
+
"model.language_model.layers.19.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
|
| 364 |
+
"model.language_model.layers.5.linear_attn.in_proj_z.weight": "model-00002-of-00004.safetensors",
|
| 365 |
+
"model.language_model.layers.4.linear_attn.norm.weight": "model-00002-of-00004.safetensors",
|
| 366 |
+
"model.visual.blocks.8.mlp.linear_fc1.weight": "model-00002-of-00004.safetensors",
|
| 367 |
+
"model.language_model.layers.15.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
|
| 368 |
+
"model.language_model.layers.22.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
|
| 369 |
+
"model.visual.blocks.21.attn.qkv.bias": "model-00002-of-00004.safetensors",
|
| 370 |
+
"model.language_model.layers.3.self_attn.k_norm.weight": "model-00002-of-00004.safetensors",
|
| 371 |
+
"model.language_model.layers.16.input_layernorm.weight": "model-00002-of-00004.safetensors",
|
| 372 |
+
"model.language_model.layers.7.input_layernorm.weight": "model-00002-of-00004.safetensors",
|
| 373 |
+
"model.language_model.layers.5.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
|
| 374 |
+
"model.visual.blocks.21.mlp.linear_fc1.weight": "model-00002-of-00004.safetensors",
|
| 375 |
+
"model.language_model.layers.28.linear_attn.in_proj_z.weight": "model-00002-of-00004.safetensors",
|
| 376 |
+
"model.language_model.layers.29.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
|
| 377 |
+
"model.visual.blocks.20.attn.qkv.weight": "model-00002-of-00004.safetensors",
|
| 378 |
+
"model.visual.blocks.20.attn.qkv.bias": "model-00002-of-00004.safetensors",
|
| 379 |
+
"model.visual.blocks.6.mlp.linear_fc1.weight": "model-00002-of-00004.safetensors",
|
| 380 |
+
"model.language_model.layers.18.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
|
| 381 |
+
"lm_head.weight": "model-00003-of-00004.safetensors",
|
| 382 |
+
"model.language_model.layers.6.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
|
| 383 |
+
"model.language_model.layers.23.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 384 |
+
"model.visual.blocks.13.attn.qkv.bias": "model-00003-of-00004.safetensors",
|
| 385 |
+
"model.language_model.layers.8.linear_attn.in_proj_a.weight": "model-00003-of-00004.safetensors",
|
| 386 |
+
"model.visual.blocks.0.mlp.linear_fc1.weight": "model-00003-of-00004.safetensors",
|
| 387 |
+
"model.language_model.layers.3.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
|
| 388 |
+
"model.language_model.layers.21.linear_attn.in_proj_qkv.weight": "model-00003-of-00004.safetensors",
|
| 389 |
+
"model.visual.blocks.21.mlp.linear_fc2.weight": "model-00003-of-00004.safetensors",
|
| 390 |
+
"model.language_model.layers.3.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 391 |
+
"model.visual.blocks.21.norm2.bias": "model-00003-of-00004.safetensors",
|
| 392 |
+
"model.visual.blocks.12.mlp.linear_fc2.weight": "model-00003-of-00004.safetensors",
|
| 393 |
+
"model.visual.blocks.1.mlp.linear_fc1.weight": "model-00003-of-00004.safetensors",
|
| 394 |
+
"model.language_model.layers.20.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
|
| 395 |
+
"model.language_model.layers.28.input_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 396 |
+
"model.visual.blocks.17.norm2.weight": "model-00003-of-00004.safetensors",
|
| 397 |
+
"model.visual.blocks.11.attn.proj.weight": "model-00003-of-00004.safetensors",
|
| 398 |
+
"model.language_model.layers.26.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
|
| 399 |
+
"model.language_model.layers.20.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
|
| 400 |
+
"model.language_model.layers.17.linear_attn.out_proj.weight": "model-00003-of-00004.safetensors",
|
| 401 |
+
"model.language_model.layers.27.self_attn.q_norm.weight": "model-00003-of-00004.safetensors",
|
| 402 |
+
"model.language_model.layers.1.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
|
| 403 |
+
"model.visual.blocks.24.mlp.linear_fc2.weight": "model-00003-of-00004.safetensors",
|
| 404 |
+
"model.visual.blocks.26.attn.proj.bias": "model-00003-of-00004.safetensors",
|
| 405 |
+
"model.language_model.layers.4.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 406 |
+
"model.visual.blocks.26.mlp.linear_fc2.bias": "model-00003-of-00004.safetensors",
|
| 407 |
+
"model.language_model.layers.5.input_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 408 |
+
"model.visual.blocks.19.norm1.bias": "model-00003-of-00004.safetensors",
|
| 409 |
+
"model.language_model.layers.23.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
|
| 410 |
+
"model.visual.blocks.25.mlp.linear_fc2.weight": "model-00003-of-00004.safetensors",
|
| 411 |
+
"model.language_model.layers.4.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
|
| 412 |
+
"model.language_model.layers.29.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
|
| 413 |
+
"model.visual.blocks.19.mlp.linear_fc1.weight": "model-00003-of-00004.safetensors",
|
| 414 |
+
"model.language_model.layers.25.input_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 415 |
+
"model.language_model.layers.30.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
|
| 416 |
+
"model.visual.blocks.16.norm1.bias": "model-00003-of-00004.safetensors",
|
| 417 |
+
"model.language_model.layers.24.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
|
| 418 |
+
"model.visual.blocks.2.norm1.weight": "model-00003-of-00004.safetensors",
|
| 419 |
+
"model.language_model.layers.2.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
|
| 420 |
+
"model.visual.blocks.17.mlp.linear_fc1.bias": "model-00003-of-00004.safetensors",
|
| 421 |
+
"model.visual.blocks.4.mlp.linear_fc1.bias": "model-00003-of-00004.safetensors",
|
| 422 |
+
"model.language_model.layers.6.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 423 |
+
"model.visual.blocks.13.attn.qkv.weight": "model-00003-of-00004.safetensors",
|
| 424 |
+
"model.language_model.layers.5.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 425 |
+
"model.language_model.layers.20.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
|
| 426 |
+
"model.visual.blocks.6.attn.qkv.weight": "model-00003-of-00004.safetensors",
|
| 427 |
+
"model.visual.blocks.16.norm2.bias": "model-00003-of-00004.safetensors",
|
| 428 |
+
"model.visual.blocks.10.attn.qkv.weight": "model-00003-of-00004.safetensors",
|
| 429 |
+
"model.language_model.layers.18.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
|
| 430 |
+
"model.language_model.layers.28.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 431 |
+
"model.visual.blocks.10.attn.proj.bias": "model-00003-of-00004.safetensors",
|
| 432 |
+
"model.visual.blocks.11.norm2.bias": "model-00003-of-00004.safetensors",
|
| 433 |
+
"model.language_model.layers.18.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
|
| 434 |
+
"model.visual.blocks.15.mlp.linear_fc1.weight": "model-00003-of-00004.safetensors",
|
| 435 |
+
"model.language_model.layers.8.linear_attn.A_log": "model-00003-of-00004.safetensors",
|
| 436 |
+
"model.language_model.layers.16.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 437 |
+
"model.visual.blocks.20.norm2.bias": "model-00003-of-00004.safetensors",
|
| 438 |
+
"model.visual.blocks.5.attn.proj.weight": "model-00003-of-00004.safetensors",
|
| 439 |
+
"model.language_model.layers.22.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
|
| 440 |
+
"model.language_model.layers.18.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
|
| 441 |
+
"model.language_model.layers.14.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
|
| 442 |
+
"model.language_model.layers.19.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 443 |
+
"model.visual.blocks.0.norm1.bias": "model-00003-of-00004.safetensors",
|
| 444 |
+
"model.language_model.layers.2.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
|
| 445 |
+
"model.language_model.layers.18.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 446 |
+
"model.visual.blocks.11.attn.qkv.bias": "model-00003-of-00004.safetensors",
|
| 447 |
+
"model.language_model.layers.4.linear_attn.in_proj_qkv.weight": "model-00003-of-00004.safetensors",
|
| 448 |
+
"model.visual.blocks.19.attn.qkv.bias": "model-00003-of-00004.safetensors",
|
| 449 |
+
"model.visual.blocks.19.mlp.linear_fc2.bias": "model-00003-of-00004.safetensors",
|
| 450 |
+
"model.language_model.layers.16.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
|
| 451 |
+
"model.language_model.layers.0.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
|
| 452 |
+
"model.visual.blocks.11.mlp.linear_fc2.bias": "model-00003-of-00004.safetensors",
|
| 453 |
+
"model.visual.blocks.4.mlp.linear_fc2.bias": "model-00003-of-00004.safetensors",
|
| 454 |
+
"model.language_model.layers.18.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
|
| 455 |
+
"model.visual.blocks.16.attn.proj.weight": "model-00003-of-00004.safetensors",
|
| 456 |
+
"model.language_model.layers.21.linear_attn.in_proj_a.weight": "model-00003-of-00004.safetensors",
|
| 457 |
+
"model.visual.blocks.0.attn.qkv.bias": "model-00003-of-00004.safetensors",
|
| 458 |
+
"model.language_model.layers.7.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 459 |
+
"model.visual.blocks.24.mlp.linear_fc1.bias": "model-00003-of-00004.safetensors",
|
| 460 |
+
"model.language_model.layers.14.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
|
| 461 |
+
"model.visual.blocks.24.norm2.weight": "model-00003-of-00004.safetensors",
|
| 462 |
+
"model.visual.blocks.16.norm1.weight": "model-00003-of-00004.safetensors",
|
| 463 |
+
"model.visual.blocks.1.attn.proj.bias": "model-00003-of-00004.safetensors",
|
| 464 |
+
"model.language_model.layers.23.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 465 |
+
"model.language_model.layers.22.input_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 466 |
+
"model.language_model.layers.12.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 467 |
+
"model.language_model.layers.13.input_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 468 |
+
"model.language_model.layers.5.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
|
| 469 |
+
"model.language_model.layers.31.self_attn.q_norm.weight": "model-00003-of-00004.safetensors",
|
| 470 |
+
"model.visual.blocks.23.mlp.linear_fc1.bias": "model-00003-of-00004.safetensors",
|
| 471 |
+
"model.language_model.layers.24.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
|
| 472 |
+
"model.visual.blocks.3.mlp.linear_fc2.weight": "model-00003-of-00004.safetensors",
|
| 473 |
+
"model.visual.blocks.21.norm1.bias": "model-00003-of-00004.safetensors",
|
| 474 |
+
"model.language_model.layers.12.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
|
| 475 |
+
"model.visual.blocks.12.norm2.weight": "model-00003-of-00004.safetensors",
|
| 476 |
+
"model.visual.blocks.12.mlp.linear_fc1.bias": "model-00003-of-00004.safetensors",
|
| 477 |
+
"model.visual.blocks.22.attn.proj.bias": "model-00003-of-00004.safetensors",
|
| 478 |
+
"model.visual.blocks.4.attn.proj.bias": "model-00003-of-00004.safetensors",
|
| 479 |
+
"model.visual.blocks.11.mlp.linear_fc2.weight": "model-00003-of-00004.safetensors",
|
| 480 |
+
"model.language_model.layers.3.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 481 |
+
"model.visual.blocks.5.attn.proj.bias": "model-00003-of-00004.safetensors",
|
| 482 |
+
"model.language_model.layers.16.linear_attn.dt_bias": "model-00003-of-00004.safetensors",
|
| 483 |
+
"model.visual.blocks.0.mlp.linear_fc2.bias": "model-00003-of-00004.safetensors",
|
| 484 |
+
"model.language_model.layers.16.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
|
| 485 |
+
"model.language_model.layers.17.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
|
| 486 |
+
"model.visual.blocks.18.norm2.weight": "model-00003-of-00004.safetensors",
|
| 487 |
+
"model.language_model.layers.20.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
|
| 488 |
+
"model.language_model.layers.0.input_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 489 |
+
"model.language_model.layers.22.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 490 |
+
"model.language_model.layers.29.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 491 |
+
"model.visual.blocks.21.mlp.linear_fc1.bias": "model-00003-of-00004.safetensors",
|
| 492 |
+
"model.language_model.layers.7.self_attn.k_norm.weight": "model-00003-of-00004.safetensors",
|
| 493 |
+
"model.language_model.layers.22.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 494 |
+
"model.visual.blocks.25.norm1.weight": "model-00003-of-00004.safetensors",
|
| 495 |
+
"model.language_model.layers.15.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
|
| 496 |
+
"model.visual.blocks.4.attn.qkv.weight": "model-00003-of-00004.safetensors",
|
| 497 |
+
"model.language_model.layers.10.linear_attn.dt_bias": "model-00003-of-00004.safetensors",
|
| 498 |
+
"model.visual.blocks.20.mlp.linear_fc2.weight": "model-00003-of-00004.safetensors",
|
| 499 |
+
"model.language_model.layers.27.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
|
| 500 |
+
"model.language_model.layers.30.linear_attn.in_proj_qkv.weight": "model-00003-of-00004.safetensors",
|
| 501 |
+
"model.language_model.layers.24.input_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 502 |
+
"model.language_model.layers.7.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
|
| 503 |
+
"model.language_model.layers.0.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
|
| 504 |
+
"model.visual.blocks.3.attn.qkv.bias": "model-00003-of-00004.safetensors",
|
| 505 |
+
"model.visual.blocks.17.norm1.bias": "model-00003-of-00004.safetensors",
|
| 506 |
+
"model.language_model.layers.5.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
|
| 507 |
+
"model.language_model.layers.14.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 508 |
+
"model.visual.blocks.24.attn.qkv.bias": "model-00003-of-00004.safetensors",
|
| 509 |
+
"model.language_model.layers.14.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 510 |
+
"model.language_model.layers.24.linear_attn.A_log": "model-00003-of-00004.safetensors",
|
| 511 |
+
"model.language_model.layers.16.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
|
| 512 |
+
"model.visual.blocks.13.attn.proj.weight": "model-00003-of-00004.safetensors",
|
| 513 |
+
"model.language_model.layers.29.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
|
| 514 |
+
"model.visual.blocks.18.mlp.linear_fc2.weight": "model-00003-of-00004.safetensors",
|
| 515 |
+
"model.language_model.layers.10.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
|
| 516 |
+
"model.visual.blocks.25.attn.proj.weight": "model-00003-of-00004.safetensors",
|
| 517 |
+
"model.language_model.layers.27.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
|
| 518 |
+
"model.visual.blocks.0.attn.proj.weight": "model-00003-of-00004.safetensors",
|
| 519 |
+
"model.visual.blocks.7.mlp.linear_fc2.weight": "model-00003-of-00004.safetensors",
|
| 520 |
+
"model.language_model.layers.9.input_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 521 |
+
"model.language_model.layers.25.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
|
| 522 |
+
"model.language_model.layers.27.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
|
| 523 |
+
"model.visual.blocks.15.attn.proj.bias": "model-00003-of-00004.safetensors",
|
| 524 |
+
"model.language_model.layers.23.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
|
| 525 |
+
"model.visual.blocks.14.norm1.bias": "model-00003-of-00004.safetensors",
|
| 526 |
+
"model.language_model.layers.20.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
|
| 527 |
+
"model.language_model.layers.30.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 528 |
+
"model.visual.blocks.8.attn.proj.bias": "model-00003-of-00004.safetensors",
|
| 529 |
+
"model.language_model.layers.22.linear_attn.A_log": "model-00003-of-00004.safetensors",
|
| 530 |
+
"model.language_model.layers.20.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
|
| 531 |
+
"model.language_model.layers.12.linear_attn.in_proj_a.weight": "model-00003-of-00004.safetensors",
|
| 532 |
+
"model.language_model.layers.29.input_layernorm.weight": "model-00003-of-00004.safetensors",
|
| 533 |
+
"model.language_model.layers.11.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
|
| 534 |
+
"model.language_model.layers.12.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 535 |
+
"model.language_model.layers.13.linear_attn.in_proj_b.weight": "model-00004-of-00004.safetensors",
|
| 536 |
+
"model.language_model.layers.25.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
|
| 537 |
+
"model.visual.blocks.26.norm2.weight": "model-00004-of-00004.safetensors",
|
| 538 |
+
"model.visual.blocks.9.mlp.linear_fc2.weight": "model-00004-of-00004.safetensors",
|
| 539 |
+
"model.visual.blocks.6.attn.proj.bias": "model-00004-of-00004.safetensors",
|
| 540 |
+
"model.visual.blocks.19.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 541 |
+
"model.visual.blocks.5.attn.qkv.bias": "model-00004-of-00004.safetensors",
|
| 542 |
+
"model.language_model.layers.8.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 543 |
+
"model.visual.blocks.7.attn.proj.weight": "model-00004-of-00004.safetensors",
|
| 544 |
+
"model.language_model.layers.1.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 545 |
+
"model.visual.blocks.15.attn.qkv.weight": "model-00004-of-00004.safetensors",
|
| 546 |
+
"model.language_model.layers.15.self_attn.q_proj.weight": "model-00004-of-00004.safetensors",
|
| 547 |
+
"model.visual.blocks.12.norm1.weight": "model-00004-of-00004.safetensors",
|
| 548 |
+
"model.visual.blocks.25.attn.proj.bias": "model-00004-of-00004.safetensors",
|
| 549 |
+
"model.language_model.layers.30.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
|
| 550 |
+
"model.language_model.layers.22.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
|
| 551 |
+
"model.language_model.layers.2.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 552 |
+
"model.visual.blocks.16.mlp.linear_fc1.weight": "model-00004-of-00004.safetensors",
|
| 553 |
+
"model.language_model.layers.11.self_attn.k_norm.weight": "model-00004-of-00004.safetensors",
|
| 554 |
+
"model.language_model.layers.10.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
|
| 555 |
+
"model.visual.blocks.3.norm2.weight": "model-00004-of-00004.safetensors",
|
| 556 |
+
"model.visual.blocks.15.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 557 |
+
"model.visual.blocks.5.norm2.weight": "model-00004-of-00004.safetensors",
|
| 558 |
+
"model.language_model.layers.7.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 559 |
+
"model.visual.blocks.23.norm2.bias": "model-00004-of-00004.safetensors",
|
| 560 |
+
"model.language_model.layers.16.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
|
| 561 |
+
"model.language_model.layers.2.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
|
| 562 |
+
"model.language_model.layers.13.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
|
| 563 |
+
"model.visual.blocks.18.attn.proj.weight": "model-00004-of-00004.safetensors",
|
| 564 |
+
"model.language_model.layers.6.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
|
| 565 |
+
"model.visual.blocks.13.mlp.linear_fc2.bias": "model-00004-of-00004.safetensors",
|
| 566 |
+
"model.language_model.layers.8.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 567 |
+
"model.visual.blocks.7.norm2.weight": "model-00004-of-00004.safetensors",
|
| 568 |
+
"model.language_model.layers.6.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
|
| 569 |
+
"model.visual.blocks.15.norm1.weight": "model-00004-of-00004.safetensors",
|
| 570 |
+
"model.language_model.layers.5.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
|
| 571 |
+
"model.language_model.layers.17.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
|
| 572 |
+
"model.visual.blocks.2.norm2.weight": "model-00004-of-00004.safetensors",
|
| 573 |
+
"model.visual.blocks.24.mlp.linear_fc2.bias": "model-00004-of-00004.safetensors",
|
| 574 |
+
"model.language_model.layers.1.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
|
| 575 |
+
"model.visual.blocks.23.mlp.linear_fc2.weight": "model-00004-of-00004.safetensors",
|
| 576 |
+
"model.visual.blocks.23.norm1.weight": "model-00004-of-00004.safetensors",
|
| 577 |
+
"model.language_model.layers.14.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
|
| 578 |
+
"model.visual.blocks.9.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 579 |
+
"model.visual.blocks.26.attn.proj.weight": "model-00004-of-00004.safetensors",
|
| 580 |
+
"model.language_model.layers.29.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
|
| 581 |
+
"model.visual.blocks.18.attn.proj.bias": "model-00004-of-00004.safetensors",
|
| 582 |
+
"model.visual.blocks.8.attn.qkv.weight": "model-00004-of-00004.safetensors",
|
| 583 |
+
"model.language_model.layers.2.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
|
| 584 |
+
"model.language_model.layers.19.self_attn.k_norm.weight": "model-00004-of-00004.safetensors",
|
| 585 |
+
"model.language_model.layers.19.self_attn.k_proj.weight": "model-00004-of-00004.safetensors",
|
| 586 |
+
"model.visual.blocks.18.mlp.linear_fc2.bias": "model-00004-of-00004.safetensors",
|
| 587 |
+
"model.language_model.layers.31.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
|
| 588 |
+
"model.language_model.layers.16.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
|
| 589 |
+
"model.language_model.layers.31.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 590 |
+
"model.visual.blocks.2.mlp.linear_fc2.bias": "model-00004-of-00004.safetensors",
|
| 591 |
+
"model.visual.blocks.14.attn.qkv.bias": "model-00004-of-00004.safetensors",
|
| 592 |
+
"model.language_model.layers.5.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 593 |
+
"model.visual.blocks.8.attn.qkv.bias": "model-00004-of-00004.safetensors",
|
| 594 |
+
"model.language_model.layers.5.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
|
| 595 |
+
"model.language_model.layers.4.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 596 |
+
"model.language_model.layers.15.self_attn.k_norm.weight": "model-00004-of-00004.safetensors",
|
| 597 |
+
"model.language_model.layers.5.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
|
| 598 |
+
"model.language_model.layers.29.linear_attn.in_proj_b.weight": "model-00004-of-00004.safetensors",
|
| 599 |
+
"model.visual.blocks.11.mlp.linear_fc1.weight": "model-00004-of-00004.safetensors",
|
| 600 |
+
"model.visual.blocks.13.mlp.linear_fc1.weight": "model-00004-of-00004.safetensors",
|
| 601 |
+
"model.visual.blocks.0.attn.proj.bias": "model-00004-of-00004.safetensors",
|
| 602 |
+
"model.language_model.layers.8.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
|
| 603 |
+
"model.language_model.layers.16.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 604 |
+
"model.language_model.layers.14.linear_attn.A_log": "model-00004-of-00004.safetensors",
|
| 605 |
+
"model.language_model.layers.31.self_attn.o_proj.weight": "model-00004-of-00004.safetensors",
|
| 606 |
+
"model.language_model.layers.9.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
|
| 607 |
+
"model.language_model.layers.4.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
|
| 608 |
+
"model.visual.blocks.26.attn.qkv.weight": "model-00004-of-00004.safetensors",
|
| 609 |
+
"model.language_model.layers.9.linear_attn.A_log": "model-00004-of-00004.safetensors",
|
| 610 |
+
"model.language_model.layers.26.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
|
| 611 |
+
"model.language_model.layers.23.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 612 |
+
"model.visual.blocks.2.attn.proj.weight": "model-00004-of-00004.safetensors",
|
| 613 |
+
"model.language_model.layers.2.linear_attn.A_log": "model-00004-of-00004.safetensors",
|
| 614 |
+
"model.language_model.layers.0.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
|
| 615 |
+
"model.language_model.layers.6.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
|
| 616 |
+
"model.visual.blocks.6.norm1.bias": "model-00004-of-00004.safetensors",
|
| 617 |
+
"model.language_model.layers.31.input_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 618 |
+
"model.language_model.layers.1.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
|
| 619 |
+
"model.visual.blocks.26.mlp.linear_fc2.weight": "model-00004-of-00004.safetensors",
|
| 620 |
+
"model.visual.blocks.24.norm2.bias": "model-00004-of-00004.safetensors",
|
| 621 |
+
"model.visual.blocks.22.norm1.weight": "model-00004-of-00004.safetensors",
|
| 622 |
+
"model.language_model.layers.24.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
|
| 623 |
+
"model.visual.blocks.18.norm2.bias": "model-00004-of-00004.safetensors",
|
| 624 |
+
"model.visual.blocks.11.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 625 |
+
"model.language_model.layers.9.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
|
| 626 |
+
"model.language_model.layers.7.self_attn.k_proj.weight": "model-00004-of-00004.safetensors",
|
| 627 |
+
"model.language_model.layers.9.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 628 |
+
"model.visual.blocks.26.norm1.bias": "model-00004-of-00004.safetensors",
|
| 629 |
+
"model.language_model.layers.12.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
|
| 630 |
+
"model.visual.blocks.10.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 631 |
+
"model.language_model.layers.25.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
|
| 632 |
+
"model.language_model.layers.24.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
|
| 633 |
+
"model.visual.blocks.25.attn.qkv.bias": "model-00004-of-00004.safetensors",
|
| 634 |
+
"model.language_model.layers.21.linear_attn.A_log": "model-00004-of-00004.safetensors",
|
| 635 |
+
"model.language_model.layers.2.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
|
| 636 |
+
"model.language_model.layers.6.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 637 |
+
"model.language_model.layers.27.self_attn.k_norm.weight": "model-00004-of-00004.safetensors",
|
| 638 |
+
"model.language_model.layers.17.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 639 |
+
"model.visual.blocks.22.attn.proj.weight": "model-00004-of-00004.safetensors",
|
| 640 |
+
"model.visual.blocks.5.attn.qkv.weight": "model-00004-of-00004.safetensors",
|
| 641 |
+
"model.language_model.layers.23.self_attn.k_norm.weight": "model-00004-of-00004.safetensors",
|
| 642 |
+
"model.language_model.layers.10.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
|
| 643 |
+
"model.visual.blocks.1.mlp.linear_fc2.bias": "model-00004-of-00004.safetensors",
|
| 644 |
+
"model.language_model.layers.1.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 645 |
+
"model.language_model.layers.12.linear_attn.A_log": "model-00004-of-00004.safetensors",
|
| 646 |
+
"model.visual.blocks.18.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 647 |
+
"model.visual.blocks.13.norm2.weight": "model-00004-of-00004.safetensors",
|
| 648 |
+
"model.language_model.layers.7.self_attn.q_proj.weight": "model-00004-of-00004.safetensors",
|
| 649 |
+
"model.language_model.layers.21.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 650 |
+
"model.visual.blocks.24.mlp.linear_fc1.weight": "model-00004-of-00004.safetensors",
|
| 651 |
+
"model.language_model.layers.27.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
|
| 652 |
+
"model.language_model.layers.30.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 653 |
+
"model.visual.blocks.16.attn.qkv.bias": "model-00004-of-00004.safetensors",
|
| 654 |
+
"model.visual.blocks.19.attn.qkv.weight": "model-00004-of-00004.safetensors",
|
| 655 |
+
"model.language_model.layers.3.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 656 |
+
"model.visual.blocks.5.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 657 |
+
"model.visual.blocks.16.norm2.weight": "model-00004-of-00004.safetensors",
|
| 658 |
+
"model.language_model.layers.13.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
|
| 659 |
+
"model.language_model.layers.26.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
|
| 660 |
+
"model.language_model.layers.27.input_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 661 |
+
"model.language_model.layers.9.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 662 |
+
"model.visual.blocks.2.mlp.linear_fc1.weight": "model-00004-of-00004.safetensors",
|
| 663 |
+
"model.language_model.layers.28.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
|
| 664 |
+
"model.visual.merger.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 665 |
+
"model.language_model.layers.20.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 666 |
+
"model.language_model.layers.17.linear_attn.in_proj_b.weight": "model-00004-of-00004.safetensors",
|
| 667 |
+
"model.language_model.layers.17.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 668 |
+
"model.language_model.layers.13.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
|
| 669 |
+
"model.visual.blocks.19.norm2.bias": "model-00004-of-00004.safetensors",
|
| 670 |
+
"model.visual.blocks.6.norm2.weight": "model-00004-of-00004.safetensors",
|
| 671 |
+
"model.language_model.layers.30.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
|
| 672 |
+
"model.language_model.layers.19.self_attn.v_proj.weight": "model-00004-of-00004.safetensors",
|
| 673 |
+
"model.language_model.layers.25.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
|
| 674 |
+
"model.visual.blocks.12.attn.qkv.weight": "model-00004-of-00004.safetensors",
|
| 675 |
+
"model.visual.blocks.10.attn.proj.weight": "model-00004-of-00004.safetensors",
|
| 676 |
+
"model.language_model.layers.11.self_attn.o_proj.weight": "model-00004-of-00004.safetensors",
|
| 677 |
+
"model.visual.blocks.4.attn.qkv.bias": "model-00004-of-00004.safetensors",
|
| 678 |
+
"model.visual.blocks.3.norm1.bias": "model-00004-of-00004.safetensors",
|
| 679 |
+
"model.language_model.layers.31.self_attn.k_norm.weight": "model-00004-of-00004.safetensors",
|
| 680 |
+
"model.language_model.layers.25.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
|
| 681 |
+
"model.visual.blocks.4.attn.proj.weight": "model-00004-of-00004.safetensors",
|
| 682 |
+
"model.visual.blocks.9.attn.qkv.bias": "model-00004-of-00004.safetensors",
|
| 683 |
+
"model.visual.blocks.0.norm1.weight": "model-00004-of-00004.safetensors",
|
| 684 |
+
"model.language_model.layers.28.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
|
| 685 |
+
"model.language_model.layers.2.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
|
| 686 |
+
"model.language_model.layers.15.self_attn.o_proj.weight": "model-00004-of-00004.safetensors",
|
| 687 |
+
"model.language_model.layers.28.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
|
| 688 |
+
"model.visual.blocks.20.norm2.weight": "model-00004-of-00004.safetensors",
|
| 689 |
+
"model.visual.blocks.26.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 690 |
+
"model.language_model.layers.26.input_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 691 |
+
"model.language_model.layers.15.self_attn.q_norm.weight": "model-00004-of-00004.safetensors",
|
| 692 |
+
"model.language_model.layers.26.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 693 |
+
"model.language_model.layers.23.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
|
| 694 |
+
"model.language_model.layers.1.linear_attn.A_log": "model-00004-of-00004.safetensors",
|
| 695 |
+
"model.visual.blocks.23.attn.qkv.bias": "model-00004-of-00004.safetensors",
|
| 696 |
+
"model.visual.blocks.22.norm2.weight": "model-00004-of-00004.safetensors",
|
| 697 |
+
"model.language_model.layers.20.linear_attn.A_log": "model-00004-of-00004.safetensors",
|
| 698 |
+
"model.visual.blocks.1.attn.qkv.weight": "model-00004-of-00004.safetensors",
|
| 699 |
+
"model.visual.merger.norm.bias": "model-00004-of-00004.safetensors",
|
| 700 |
+
"model.language_model.layers.30.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
|
| 701 |
+
"model.visual.blocks.6.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 702 |
+
"model.language_model.layers.22.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 703 |
+
"model.visual.blocks.19.mlp.linear_fc2.weight": "model-00004-of-00004.safetensors",
|
| 704 |
+
"model.language_model.layers.2.input_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 705 |
+
"model.visual.blocks.11.attn.proj.bias": "model-00004-of-00004.safetensors",
|
| 706 |
+
"model.visual.blocks.9.norm2.bias": "model-00004-of-00004.safetensors",
|
| 707 |
+
"model.language_model.layers.29.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
|
| 708 |
+
"model.visual.blocks.23.norm1.bias": "model-00004-of-00004.safetensors",
|
| 709 |
+
"model.language_model.layers.12.input_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 710 |
+
"model.language_model.layers.24.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 711 |
+
"model.visual.blocks.23.mlp.linear_fc1.weight": "model-00004-of-00004.safetensors",
|
| 712 |
+
"model.language_model.layers.25.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
|
| 713 |
+
"model.language_model.layers.20.input_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 714 |
+
"model.language_model.layers.27.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 715 |
+
"model.language_model.layers.28.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 716 |
+
"model.language_model.layers.1.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
|
| 717 |
+
"model.visual.blocks.14.mlp.linear_fc1.weight": "model-00004-of-00004.safetensors",
|
| 718 |
+
"model.language_model.layers.26.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
|
| 719 |
+
"model.language_model.layers.8.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
|
| 720 |
+
"model.language_model.layers.4.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
|
| 721 |
+
"model.language_model.layers.0.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 722 |
+
"model.visual.blocks.9.norm1.bias": "model-00004-of-00004.safetensors",
|
| 723 |
+
"model.language_model.layers.17.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
|
| 724 |
+
"model.language_model.layers.0.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
|
| 725 |
+
"model.visual.blocks.1.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 726 |
+
"model.language_model.layers.11.self_attn.q_norm.weight": "model-00004-of-00004.safetensors",
|
| 727 |
+
"model.visual.blocks.2.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 728 |
+
"model.language_model.layers.10.input_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 729 |
+
"model.visual.blocks.17.mlp.linear_fc2.bias": "model-00004-of-00004.safetensors",
|
| 730 |
+
"model.visual.blocks.21.attn.proj.weight": "model-00004-of-00004.safetensors",
|
| 731 |
+
"model.visual.blocks.7.norm2.bias": "model-00004-of-00004.safetensors",
|
| 732 |
+
"model.visual.blocks.20.norm1.weight": "model-00004-of-00004.safetensors",
|
| 733 |
+
"model.language_model.norm.weight": "model-00004-of-00004.safetensors",
|
| 734 |
+
"model.language_model.layers.4.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
|
| 735 |
+
"model.visual.blocks.15.mlp.linear_fc2.weight": "model-00004-of-00004.safetensors",
|
| 736 |
+
"model.language_model.layers.29.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 737 |
+
"model.visual.blocks.17.mlp.linear_fc1.weight": "model-00004-of-00004.safetensors",
|
| 738 |
+
"model.visual.blocks.24.attn.proj.weight": "model-00004-of-00004.safetensors",
|
| 739 |
+
"model.language_model.layers.26.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
|
| 740 |
+
"model.visual.blocks.19.norm1.weight": "model-00004-of-00004.safetensors",
|
| 741 |
+
"model.visual.blocks.22.norm1.bias": "model-00004-of-00004.safetensors",
|
| 742 |
+
"model.language_model.layers.10.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
|
| 743 |
+
"model.visual.blocks.7.norm1.weight": "model-00004-of-00004.safetensors",
|
| 744 |
+
"model.language_model.layers.17.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
|
| 745 |
+
"model.visual.blocks.1.attn.proj.weight": "model-00004-of-00004.safetensors",
|
| 746 |
+
"model.visual.blocks.8.mlp.linear_fc1.bias": "model-00004-of-00004.safetensors",
|
| 747 |
+
"model.language_model.layers.30.input_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 748 |
+
"model.language_model.layers.21.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 749 |
+
"model.visual.blocks.10.norm2.bias": "model-00004-of-00004.safetensors",
|
| 750 |
+
"model.visual.blocks.7.mlp.linear_fc1.weight": "model-00004-of-00004.safetensors",
|
| 751 |
+
"model.language_model.layers.14.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
|
| 752 |
+
"model.visual.blocks.1.attn.qkv.bias": "model-00004-of-00004.safetensors",
|
| 753 |
+
"model.language_model.layers.29.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
|
| 754 |
+
"model.visual.blocks.13.norm2.bias": "model-00004-of-00004.safetensors",
|
| 755 |
+
"model.visual.blocks.18.norm1.bias": "model-00004-of-00004.safetensors",
|
| 756 |
+
"model.language_model.layers.17.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
|
| 757 |
+
"model.visual.blocks.11.norm1.bias": "model-00004-of-00004.safetensors",
|
| 758 |
+
"model.language_model.layers.14.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
|
| 759 |
+
"model.language_model.layers.8.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
|
| 760 |
+
"model.language_model.layers.11.self_attn.q_proj.weight": "model-00004-of-00004.safetensors",
|
| 761 |
+
"model.language_model.layers.21.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
|
| 762 |
+
"model.language_model.layers.12.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
|
| 763 |
+
"model.visual.blocks.6.norm2.bias": "model-00004-of-00004.safetensors",
|
| 764 |
+
"model.language_model.layers.14.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
|
| 765 |
+
"model.visual.blocks.0.mlp.linear_fc2.weight": "model-00004-of-00004.safetensors"
|
| 766 |
+
}
|
| 767 |
+
}
|
preprocessor_config.json
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"do_convert_rgb": true,
|
| 3 |
+
"do_normalize": true,
|
| 4 |
+
"do_rescale": true,
|
| 5 |
+
"do_resize": true,
|
| 6 |
+
"image_mean": [
|
| 7 |
+
0.5,
|
| 8 |
+
0.5,
|
| 9 |
+
0.5
|
| 10 |
+
],
|
| 11 |
+
"image_processor_type": "Qwen2VLImageProcessor",
|
| 12 |
+
"image_std": [
|
| 13 |
+
0.5,
|
| 14 |
+
0.5,
|
| 15 |
+
0.5
|
| 16 |
+
],
|
| 17 |
+
"merge_size": 2,
|
| 18 |
+
"patch_size": 16,
|
| 19 |
+
"resample": 3,
|
| 20 |
+
"rescale_factor": 0.00392156862745098,
|
| 21 |
+
"size": {
|
| 22 |
+
"longest_edge": 16777216,
|
| 23 |
+
"shortest_edge": 65536
|
| 24 |
+
},
|
| 25 |
+
"temporal_patch_size": 2
|
| 26 |
+
}
|
processor_config.json
ADDED
|
@@ -0,0 +1,60 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"image_processor": {
|
| 3 |
+
"do_convert_rgb": true,
|
| 4 |
+
"do_normalize": true,
|
| 5 |
+
"do_rescale": true,
|
| 6 |
+
"do_resize": true,
|
| 7 |
+
"image_mean": [
|
| 8 |
+
0.5,
|
| 9 |
+
0.5,
|
| 10 |
+
0.5
|
| 11 |
+
],
|
| 12 |
+
"image_processor_type": "Qwen2VLImageProcessor",
|
| 13 |
+
"image_std": [
|
| 14 |
+
0.5,
|
| 15 |
+
0.5,
|
| 16 |
+
0.5
|
| 17 |
+
],
|
| 18 |
+
"merge_size": 2,
|
| 19 |
+
"patch_size": 16,
|
| 20 |
+
"resample": 3,
|
| 21 |
+
"rescale_factor": 0.00392156862745098,
|
| 22 |
+
"size": {
|
| 23 |
+
"longest_edge": 16777216,
|
| 24 |
+
"shortest_edge": 65536
|
| 25 |
+
},
|
| 26 |
+
"temporal_patch_size": 2
|
| 27 |
+
},
|
| 28 |
+
"processor_class": "Qwen3VLProcessor",
|
| 29 |
+
"video_processor": {
|
| 30 |
+
"do_convert_rgb": true,
|
| 31 |
+
"do_normalize": true,
|
| 32 |
+
"do_rescale": true,
|
| 33 |
+
"do_resize": true,
|
| 34 |
+
"do_sample_frames": true,
|
| 35 |
+
"fps": 2,
|
| 36 |
+
"image_mean": [
|
| 37 |
+
0.5,
|
| 38 |
+
0.5,
|
| 39 |
+
0.5
|
| 40 |
+
],
|
| 41 |
+
"image_std": [
|
| 42 |
+
0.5,
|
| 43 |
+
0.5,
|
| 44 |
+
0.5
|
| 45 |
+
],
|
| 46 |
+
"max_frames": 768,
|
| 47 |
+
"merge_size": 2,
|
| 48 |
+
"min_frames": 4,
|
| 49 |
+
"patch_size": 16,
|
| 50 |
+
"resample": 3,
|
| 51 |
+
"rescale_factor": 0.00392156862745098,
|
| 52 |
+
"return_metadata": false,
|
| 53 |
+
"size": {
|
| 54 |
+
"longest_edge": 25165824,
|
| 55 |
+
"shortest_edge": 4096
|
| 56 |
+
},
|
| 57 |
+
"temporal_patch_size": 2,
|
| 58 |
+
"video_processor_type": "Qwen3VLVideoProcessor"
|
| 59 |
+
}
|
| 60 |
+
}
|
serve.py
ADDED
|
@@ -0,0 +1,167 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""blink server: a Jev-compatible decision endpoint.
|
| 2 |
+
|
| 3 |
+
pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" accelerate safetensors huggingface_hub
|
| 4 |
+
hf download thegovind/blink-4b --revision v1.0 --local-dir blink-4b
|
| 5 |
+
python blink-4b/serve.py --model ./blink-4b --port 8000
|
| 6 |
+
|
| 7 |
+
POST /v1/systemone {"state": ..., "questions": {...}} -> {"model", "answers", "usage"}
|
| 8 |
+
GET /healthz -> {"ok", "model", "revision", "weights_verified", "hub_offline", "warmup", "kernels", "versions"}
|
| 9 |
+
|
| 10 |
+
Requests are served one at a time. A request over a limit (options per choice, context length,
|
| 11 |
+
questions per request) gets HTTP 422 with the reason; nothing is truncated. With weights.sha256 beside
|
| 12 |
+
the weights, every listed file is hashed before serving (weights_verified). Serving a local folder switches
|
| 13 |
+
the Hugging Face libraries to offline mode before any of them loads (hub_offline reports the setting the
|
| 14 |
+
libraries actually use); the server does not otherwise restrict the network.
|
| 15 |
+
"""
|
| 16 |
+
|
| 17 |
+
from __future__ import annotations
|
| 18 |
+
|
| 19 |
+
import argparse
|
| 20 |
+
import hashlib
|
| 21 |
+
import json
|
| 22 |
+
import os
|
| 23 |
+
import socket
|
| 24 |
+
import sys
|
| 25 |
+
import threading
|
| 26 |
+
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
| 27 |
+
|
| 28 |
+
WARMUP = ("Order 4471 arrived with a cracked screen. The customer attached photos and wants a replacement.",
|
| 29 |
+
{"route": {"type": "choice", "instructions": "Which team should handle this?",
|
| 30 |
+
"criteria": {"returns": "Damaged or wrong items", "billing": "Charges and refunds",
|
| 31 |
+
"shipping": "Late or lost parcels"}},
|
| 32 |
+
"urgent": {"type": "noul", "instructions": "Does this need a reply today?"}})
|
| 33 |
+
|
| 34 |
+
|
| 35 |
+
def sha256(path: str) -> str:
|
| 36 |
+
h = hashlib.sha256()
|
| 37 |
+
with open(path, "rb") as fh:
|
| 38 |
+
for block in iter(lambda: fh.read(1 << 24), b""):
|
| 39 |
+
h.update(block)
|
| 40 |
+
return h.hexdigest()
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
def verify(root: str):
|
| 44 |
+
"""True/False against weights.sha256 ("<sha256> <file>" lines) beside the weights; None without one."""
|
| 45 |
+
manifest = os.path.join(root, "weights.sha256")
|
| 46 |
+
if not os.path.exists(manifest):
|
| 47 |
+
return None, []
|
| 48 |
+
bad = []
|
| 49 |
+
with open(manifest, encoding="utf-8") as fh:
|
| 50 |
+
for line in fh:
|
| 51 |
+
if line.strip():
|
| 52 |
+
digest, name = line.split(None, 1)
|
| 53 |
+
name = name.strip()
|
| 54 |
+
path = os.path.join(root, name)
|
| 55 |
+
if not os.path.exists(path) or sha256(path) != digest:
|
| 56 |
+
bad.append(name)
|
| 57 |
+
return not bad, bad
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
def versions() -> dict:
|
| 61 |
+
out = {}
|
| 62 |
+
for mod in ("torch", "transformers", "fla"):
|
| 63 |
+
try:
|
| 64 |
+
out[mod] = __import__(mod).__version__
|
| 65 |
+
except Exception:
|
| 66 |
+
out[mod] = None
|
| 67 |
+
return out
|
| 68 |
+
|
| 69 |
+
|
| 70 |
+
def main() -> None:
|
| 71 |
+
ap = argparse.ArgumentParser(description="Serve blink over a Jev-compatible HTTP API.")
|
| 72 |
+
ap.add_argument("--model", default=os.environ.get("BLINK_MODEL", "thegovind/blink-4b"))
|
| 73 |
+
ap.add_argument("--revision", default=os.environ.get("BLINK_REVISION"))
|
| 74 |
+
ap.add_argument("--host", default="127.0.0.1")
|
| 75 |
+
ap.add_argument("--port", type=int, default=8000)
|
| 76 |
+
a = ap.parse_args()
|
| 77 |
+
|
| 78 |
+
local = os.path.isdir(a.model)
|
| 79 |
+
if local:
|
| 80 |
+
# the Hub libraries read these once, when they are first imported, so they must be set before that
|
| 81 |
+
for var in ("HF_HUB_OFFLINE", "TRANSFORMERS_OFFLINE", "HF_HUB_DISABLE_TELEMETRY"):
|
| 82 |
+
os.environ.setdefault(var, "1")
|
| 83 |
+
os.environ["BLINK_ENGINE"] = "torch"
|
| 84 |
+
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
| 85 |
+
import blink
|
| 86 |
+
|
| 87 |
+
if local:
|
| 88 |
+
root = a.model
|
| 89 |
+
else:
|
| 90 |
+
from huggingface_hub import snapshot_download
|
| 91 |
+
|
| 92 |
+
root = snapshot_download(a.model, revision=a.revision)
|
| 93 |
+
verified, bad = verify(root)
|
| 94 |
+
if verified is False:
|
| 95 |
+
sys.exit(f"weights.sha256 mismatch: {', '.join(bad)}")
|
| 96 |
+
engine = blink.TorchEngine(root, None, blink.TEMPERATURE)
|
| 97 |
+
blink._ENGINE = engine
|
| 98 |
+
first = blink.decide(*WARMUP)["answers"]
|
| 99 |
+
repeat_identical = blink.decide(*WARMUP)["answers"] == first
|
| 100 |
+
ver = versions()
|
| 101 |
+
kernels = f"flash-linear-attention {ver['fla']}" if ver["fla"] else "reference (much slower; install flash-linear-attention)"
|
| 102 |
+
try:
|
| 103 |
+
from huggingface_hub import constants as hub_constants
|
| 104 |
+
|
| 105 |
+
hub_offline = bool(hub_constants.HF_HUB_OFFLINE)
|
| 106 |
+
except Exception:
|
| 107 |
+
hub_offline = None
|
| 108 |
+
health = {"ok": True, "model": a.model, "revision": a.revision, "weights_verified": verified,
|
| 109 |
+
"hub_offline": hub_offline, "warmup": {"repeat_identical": repeat_identical}, "kernels": kernels,
|
| 110 |
+
"versions": ver}
|
| 111 |
+
lock = threading.Lock()
|
| 112 |
+
|
| 113 |
+
class Handler(BaseHTTPRequestHandler):
|
| 114 |
+
protocol_version = "HTTP/1.1"
|
| 115 |
+
|
| 116 |
+
def log_message(self, fmt, *args): # quiet by default
|
| 117 |
+
pass
|
| 118 |
+
|
| 119 |
+
def setup(self):
|
| 120 |
+
super().setup()
|
| 121 |
+
# a response is two writes (headers, then body); without TCP_NODELAY the body waits
|
| 122 |
+
# on the client's delayed ACK, a flat ~40 ms on every request of a kept-alive connection
|
| 123 |
+
self.connection.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
|
| 124 |
+
|
| 125 |
+
def _send(self, code: int, obj: dict) -> None:
|
| 126 |
+
body = json.dumps(obj, ensure_ascii=False).encode("utf-8")
|
| 127 |
+
self.send_response(code)
|
| 128 |
+
self.send_header("Content-Type", "application/json")
|
| 129 |
+
self.send_header("Content-Length", str(len(body)))
|
| 130 |
+
self.end_headers()
|
| 131 |
+
self.wfile.write(body)
|
| 132 |
+
|
| 133 |
+
def do_GET(self):
|
| 134 |
+
if self.path.rstrip("/") in ("/healthz", "/health"):
|
| 135 |
+
return self._send(200, health)
|
| 136 |
+
return self._send(404, {"error": "not found"})
|
| 137 |
+
|
| 138 |
+
def do_POST(self):
|
| 139 |
+
if self.path.rstrip("/") != "/v1/systemone":
|
| 140 |
+
return self._send(404, {"error": "not found"})
|
| 141 |
+
try:
|
| 142 |
+
size = int(self.headers.get("Content-Length") or 0)
|
| 143 |
+
req = json.loads(self.rfile.read(size) or b"{}")
|
| 144 |
+
except (ValueError, json.JSONDecodeError) as exc:
|
| 145 |
+
return self._send(400, {"error": f"invalid JSON: {exc}"})
|
| 146 |
+
if not isinstance(req, dict):
|
| 147 |
+
return self._send(400, {"error": "the body must be a JSON object"})
|
| 148 |
+
try:
|
| 149 |
+
with lock:
|
| 150 |
+
out = blink.decide(req.get("state"), req.get("questions"))
|
| 151 |
+
except blink.BlinkError as exc:
|
| 152 |
+
return self._send(422, {"error": str(exc)})
|
| 153 |
+
except Exception as exc: # noqa: BLE001 - report, keep serving
|
| 154 |
+
return self._send(500, {"error": f"{type(exc).__name__}: {exc}"})
|
| 155 |
+
return self._send(200, {
|
| 156 |
+
"model": a.model,
|
| 157 |
+
"answers": out["answers"],
|
| 158 |
+
"usage": {"input_tokens": out["meta"]["input_tokens"], "output_tokens": 0},
|
| 159 |
+
})
|
| 160 |
+
|
| 161 |
+
server = ThreadingHTTPServer((a.host, a.port), Handler)
|
| 162 |
+
print(f"blink serving {a.model} on http://{a.host}:{a.port} ({kernels}; weights_verified={verified})", flush=True)
|
| 163 |
+
server.serve_forever()
|
| 164 |
+
|
| 165 |
+
|
| 166 |
+
if __name__ == "__main__":
|
| 167 |
+
main()
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523
|
| 3 |
+
size 19989325
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"audio_bos_token": "<|audio_start|>",
|
| 4 |
+
"audio_eos_token": "<|audio_end|>",
|
| 5 |
+
"audio_token": "<|audio_pad|>",
|
| 6 |
+
"backend": "tokenizers",
|
| 7 |
+
"bos_token": null,
|
| 8 |
+
"clean_up_tokenization_spaces": false,
|
| 9 |
+
"eos_token": "<|im_end|>",
|
| 10 |
+
"errors": "replace",
|
| 11 |
+
"image_token": "<|image_pad|>",
|
| 12 |
+
"is_local": true,
|
| 13 |
+
"local_files_only": false,
|
| 14 |
+
"model_max_length": 262144,
|
| 15 |
+
"model_specific_special_tokens": {
|
| 16 |
+
"audio_bos_token": "<|audio_start|>",
|
| 17 |
+
"audio_eos_token": "<|audio_end|>",
|
| 18 |
+
"audio_token": "<|audio_pad|>",
|
| 19 |
+
"image_token": "<|image_pad|>",
|
| 20 |
+
"video_token": "<|video_pad|>",
|
| 21 |
+
"vision_bos_token": "<|vision_start|>",
|
| 22 |
+
"vision_eos_token": "<|vision_end|>"
|
| 23 |
+
},
|
| 24 |
+
"pad_token": "<|endoftext|>",
|
| 25 |
+
"pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
|
| 26 |
+
"processor_class": "Qwen3VLProcessor",
|
| 27 |
+
"split_special_tokens": false,
|
| 28 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 29 |
+
"unk_token": null,
|
| 30 |
+
"video_token": "<|video_pad|>",
|
| 31 |
+
"vision_bos_token": "<|vision_start|>",
|
| 32 |
+
"vision_eos_token": "<|vision_end|>"
|
| 33 |
+
}
|
video_preprocessor_config.json
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"size": {
|
| 3 |
+
"longest_edge": 25165824,
|
| 4 |
+
"shortest_edge": 4096
|
| 5 |
+
},
|
| 6 |
+
"patch_size": 16,
|
| 7 |
+
"temporal_patch_size": 2,
|
| 8 |
+
"merge_size": 2,
|
| 9 |
+
"image_mean": [
|
| 10 |
+
0.5,
|
| 11 |
+
0.5,
|
| 12 |
+
0.5
|
| 13 |
+
],
|
| 14 |
+
"image_std": [
|
| 15 |
+
0.5,
|
| 16 |
+
0.5,
|
| 17 |
+
0.5
|
| 18 |
+
],
|
| 19 |
+
"processor_class": "Qwen3VLProcessor",
|
| 20 |
+
"video_processor_type": "Qwen3VLVideoProcessor"
|
| 21 |
+
}
|
vocab.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
weights.sha256
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
59a64ebb4df6d1489d09a91267cf3ceb106162d4a893c4f84833cfb8c897ff63 chat_template.jinja
|
| 2 |
+
407c46388b8fa2ae9bf69fe27d40af236d373e86a6b48f5284c86df5cd183633 config.json
|
| 3 |
+
aaed5b26f1cae55c1ceb58fc483c3cd65ee8d386b61daec3d7a9df82432d9205 generation_config.json
|
| 4 |
+
a9d356d7bdf1ef4949e3e748e95b8e10ad9d4e2e838eddc38a0a7b6b94d1db8d merges.txt
|
| 5 |
+
9da483e4c921f161562d994709fe7181a05d1a14692026d9943da80021247d82 model-00001-of-00004.safetensors
|
| 6 |
+
efbf9af22c3f00f32289c50f9cd9ed5abc6cf9ab4af957d9adae1c80cd0daeae model-00002-of-00004.safetensors
|
| 7 |
+
3fb71b091963d851c95352c3d3a24ecefc114a3b281e916943a129f173db05d7 model-00003-of-00004.safetensors
|
| 8 |
+
6c8a6365689c40acdf1c7605819323cd09f7f5f8818497e9cd56f50b7ef38d34 model-00004-of-00004.safetensors
|
| 9 |
+
39fd8a226de2d7e132ef77547c3ec772368f329718da4d4808d4dd7dbe4fbd2e model.safetensors.index.json
|
| 10 |
+
3a159dfec9978a186a72ba085e0ad6a050f3968d8b364218d7bd13f5c89381f2 preprocessor_config.json
|
| 11 |
+
d89ef49ce9cd37fbf510158e13c1ef063d9286411c1ec9049932dbe0487143b1 processor_config.json
|
| 12 |
+
06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523 tokenizer.json
|
| 13 |
+
792fa3f0cb88b111e54ef3134c873531008c4df471d108da17903426e308aa7b tokenizer_config.json
|
| 14 |
+
7768af27c1fafa9cc9011c1dc20067e03f8915e03b63504550e11d5066986d13 video_preprocessor_config.json
|
| 15 |
+
ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003 vocab.json
|