---
license: other
license_name: blink-research
license_link: LICENSE.md
base_model: XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
library_name: transformers
pipeline_tag: image-text-to-text
inference: false
tags:
- decision-model
- typed-decisions
- one-pass
- option-probabilities
language:
- en
---
# blink-mimo-9b

Pick the next click from a screenshot. blink-mimo-9b reads it with its own vision tower and scores the offered targets as a typed choice, returning probabilities instead of generated text. It also takes text or JSON `state` with `choice`, `noul` (yes/no), or `score` questions. Each batch is one forward pass; large requests may need several batches.
**Try it:** [Screen click](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen) · [Space](https://huggingface.co/spaces/thegovind/blink) · [Computer use](https://thegovind.github.io/blink/computer-use/) · [API](https://thegovind.github.io/blink/api/) · [Docs](https://thegovind.github.io/blink/) · [GitHub](https://github.com/thegovind/blink) · [On a Mac (MLX)](#run-it-on-a-mac-mlx) · [blink-4b](https://huggingface.co/thegovind/blink-4b) · [blink-27b](https://huggingface.co/thegovind/blink-27b)
## At a glance
| Attribute | Detail |
|---|---|
| Base model | [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) |
| Weights size | 18.8 GB bf16 (8.95B text parameters, 9.41B total) |
| Revision | v1.4 (code revision; weights identical to v1.0) |
| License | Weights: non-commercial research and evaluation only ([LICENSE.md](LICENSE.md)); code: Apache-2.0. |
## Computer use
Screen click uses this model to score numbered targets in a screenshot as a typed choice.
### How it reads a screenshot
MiMo's own vision encoder supplies image tokens; blink scores the offered clicks in a forward pass without generating text.
From pixels to probabilities
- A screenshot with numbered boxes goes through blink-mimo-9b's own vision encoder from [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B). It has 27 encoder blocks, hidden 1152, and 333 vision tensors unchanged by training.
- The encoder turns the screenshot into image tokens.
- The language model reads the image tokens plus the typed `choice` question and offered targets.
- blink reads next-token logits for the offered options and applies an FP32 softmax. No text is generated.
- Enable self-hosted screenshot input with `--vision`.
## Computer use in ten apps
blink-mimo-9b finished 133/230 tasks (58%) zero-shot. It's the model behind [Screen click](https://huggingface.co/spaces/thegovind/blink?tab=computer-use&shot=catalog), which runs live in the Space. It uses its own vision encoder with `--vision`, not a borrowed tower. In the "Teach it your app" v1 experiment, 150 labelled screens lifted success on practised apps from 96/160 to 132/160 (60% to 83%). But it paused on only 35/76 truly risky clicks there, down from 62/62 before training. Apps it never saw went from 35/60 to 49/60. Screenshot training is still an experiment, not in the public trainer. These are original apps built for this showcase, not a public benchmark. [Watch the runs](https://huggingface.co/spaces/thegovind/blink?tab=computer-use) · [Videos and traces](https://huggingface.co/datasets/thegovind/blink-cua) · [Code](https://github.com/thegovind/blink/tree/main/examples/cua).
[](https://huggingface.co/spaces/thegovind/blink?tab=computer-use&scenario=desktop&model=blink-mimo-9b)
| App | Tasks done |
|---|---:|
| Overall | 133/230 (58%) |
| Settings | 30/30 |
| Desktop (canvas) | 15/20 |
| Files | 22/30 |
## Screenshots (opt-in, self-hosted)
Self-hosted image input is off by default: start `serve.py` with `--vision` to use MiMo's unchanged tower (image mode needs `torchvision==0.28.0`). TypeSafe's hosted Jev is text-only; screenshots are a blink self-hosted extension, not a hosted Jev feature.
Send a `data:image/png;base64,...` URI inside a string in `state`, or data URIs in a top-level `images` list, with a `choice` question whose criteria name the clickable targets. JPEG and WebP data URIs also work. Image URLs are never fetched.
These are **development readouts, not a benchmark**. With the default `first` image layout and five offered targets, target accuracy was:
| Development set | K5 target accuracy |
|---|---:|
| ScreenSpot-v2 | 94.10% |
| GUIOdyssey labelled subset | 78.95% |
Target accuracy is not completed-task success. The default layout was picked on previously seen development data. Try [catalog](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen&shot=catalog), [directory](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen&shot=directory), [lookup](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen&shot=lookup), or [done](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen&shot=done); [Next click (text)](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=nextclick) shows the page-elements-as-text path.
For text-only browser agents using page elements as text, [blink-27b](https://huggingface.co/thegovind/blink-27b) did best in local runs; see [browser-agent setup](https://thegovind.github.io/blink/api/#browser-agents). These local runs are not a benchmark.
## Quickstart
This example uses text. For screenshots, use the opt-in server path above.
```python
# pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
import os, sys
from huggingface_hub import hf_hub_download
os.environ["BLINK_MODEL"] = "thegovind/blink-mimo-9b"
os.environ["BLINK_REVISION"] = "v1.4"
sys.path.insert(0, os.path.dirname(hf_hub_download("thegovind/blink-mimo-9b", "blink.py", revision="v1.4")))
import blink
out = blink.decide(
"Order #4411 arrived with a cracked screen. I want my money back, not another one.",
{
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": {"refund": "Money back", "replacement": "A new unit", "info": "Information only"},
},
"urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
"anger": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]},
},
)
print(out["answers"]["intent"]["probabilities"])
```
## Run it as a server
`serve.py` implements TypeSafe's text decision API at `POST /v1/systemone` and `GET /v1/models`. Point server-side Python or JavaScript SDKs at it with `TYPESAFE_BASE_URL`. It processes requests one at a time by default; use `--batch-window-ms 5` for cross-request batching. `GET /healthz` reports readiness. The command below starts text-only; pass `--vision` for screenshots.
```sh
pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
hf download thegovind/blink-mimo-9b --revision v1.4 --local-dir blink-mimo-9b
python blink-mimo-9b/serve.py --model ./blink-mimo-9b --port 8000
# TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any
```
Or use Docker from the downloaded folder:
```sh
cd blink-mimo-9b
docker build -t blink-mimo-9b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-mimo-9b
```
Earlier vLLM checks passed, but the final read matched earlier answers and still failed hard ECE (0.122 vs 0.121). No MiMo opt-in server ships; use `serve.py`.
## Run it on a Mac (MLX)
[`blink_mlx.py`](https://github.com/thegovind/blink/tree/main/examples/mlx) keeps this repo's `blink.py` for prompts, option labels and answers, and runs the
forward pass with MLX on Apple silicon. Text and JSON state only, bf16 weights only; it is not a server. It loads only the text side of this checkpoint.
```sh
pip install "mlx-lm>=0.31.3"
curl -O https://raw.githubusercontent.com/thegovind/blink/main/examples/mlx/blink_mlx.py
python blink_mlx.py --model thegovind/blink-mimo-9b
```
`--check` compares your Mac's answers with the Space's saved runs. The [MLX guide](https://github.com/thegovind/blink/tree/main/examples/mlx) covers memory and limits.
## Results
| Local development readout | Result |
|---|---:|
| Decision Index 0.2 balanced skill | 43.36 |
| JevBench public hard items | 77/111; hard ECE 0.136 |
No official JevBench score for blink has been published. The JevBench numbers are public-item development proxies, not official scores or claims of rank or parity. Decision Index is a descriptive local run of the official kit, not a leaderboard submission; public training exposure affects its scores.
Model details: architecture, training, data
### Architecture and readout


MiMo's text side has 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 4096, untied embeddings, and 8.95B text parameters. The tower has 27 encoder blocks, hidden 1152; all 333 vision tensors unchanged. Fine-tuning merged 43.3M LoRA parameters at rank 16, alpha 32; 179 other language-model tensors stayed unchanged. Targets: `q_proj`, `k_proj`, `v_proj`, `o_proj`; `in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `out_proj`; `gate_proj`, `up_proj`, `down_proj`.
Readout uses FP32 softmax on next-token logits for offered labels only. Probabilities are conditional on the options, not certified chances of success.
### Training

One supervised epoch: 123,195 question rows, lr 5e-5, 615 steps. Cross-entropy trained on target distributions; no RL or preference optimization.
| Stage | Question rows | Mix |
|---|---:|---|
| MiMo | 123,195 | 61,394 public-source · 23,894 program-generated reasoning · 12,000 decision worlds · 7,860 teacher-written rows · 7,000 judge-style · 6,000 chess move choices · 5,047 exact-probability worlds |
### Data sources and licences
| Source | Licence |
|---|---|
| MMLU auxiliary train, CommonsenseQA, GSM8K | MIT |
| AQuA-RAT, Amazon ESCI | Apache-2.0 |
| searchless_chess | data CC BY 4.0 (Lichess-derived portions CC0); code Apache-2.0 |
| MedMCQA | Apache-2.0 (dataset card) |
| SuperGPQA | ODC-BY |
| WANLI, ContractNLI, BANKING77 | CC BY 4.0 |
| ARC | CC BY-SA 4.0 |
| BoolQ, Dolly-15k | CC BY-SA 3.0 |
| ANLI | CC BY-NC 4.0 |
| SciQ | CC BY-NC 3.0 |
| iSarcasmEval | MIT (upstream repository licence) |
| VAST, Humicroedit, OpenBookQA | None stated by source |
| Code-generated worlds and teacher-written documents (Qwen3.8-27B) | See LICENSE.md |
Source-repository licences do not settle rights in underlying texts.
### Evaluation and limits
The DI-S selection sample scored 55.3; the archived Decision Index 0.1 local run scored 56.53. No MMLU-Pro or GPQA items were direct training sources, though SuperGPQA rows matched some added-request text. Public train splits overlap Decision Index, and semantic or pretraining overlap cannot be ruled out. The v1.2 `blink.py` and `serve.py` used its text side only. The screenshot readouts above are development measurements, not final held-out image evaluations.
Use `--image-layout first|inline` to place image tokens (`first` is the default), or `--model-name blink-mimo-9b` for a renamed folder. Neither flag switches on images. Text limits are 255 options per choice and 2–10 score levels; invalid images or over-limit requests return 422. English-centric; does not chat or explain answers, and long policies and date/number reasoning remain weak spots.
## License
Code: Apache-2.0. Weights: non-commercial research and evaluation only; see each model card's license.
See [LICENSE.md](LICENSE.md) for the weight terms. The MiMo model card declares MIT without a separate upstream licence file ([LICENSE-MiMo.md](LICENSE-MiMo.md)); its Qwen base is Apache-2.0 (`LICENSE-Qwen`).