--- license: other license_name: blink-research license_link: LICENSE.md base_model: XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B library_name: transformers pipeline_tag: image-text-to-text inference: false tags: - decision-model - typed-decisions - one-pass - option-probabilities language: - en --- # blink-mimo-9b ![blink-mimo-9b Screen click diagram: the “Open the result” preset, a screenshot of a book catalog page with 3 numbered boxes: box 1 on the select “Book category: Maps”, box 2 on the button “Search catalog” and box 3 on the link “Open Atlas Field Guide”. The task is “Search the Maps category, then open Atlas Field Guide from the results.” The screenshot is resized to multiples of 32 px (1440 × 864) and read by blink-mimo-9b's own vision encoder from XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B (27 blocks, hidden 1152 → 4096; its 333 tensors unchanged by blink training), which turns it into 1,215 image tokens (45 × 27, one per 32 × 32 px). Three typed questions are read with the image tokens, one forward pass each, with no text generated: which box to act on next (a choice over boxes 1 to 3), is the task already done (yes or no), and would the next click be risky (yes or no). In this saved run, as the Space shows it (saved run · blink-mimo-9b), box 3 has the top probability, 62%, then box 2 33% and box 1 5%; done is 19% yes and risky 2% yes, so the verdict is “Click 3”. Demo only; it never clicks.](assets/blink-mimo-9b-screens.png) Pick the next click from a screenshot. blink-mimo-9b reads it with its own vision tower and scores the offered targets as a typed choice, returning probabilities instead of generated text. It also takes text or JSON `state` with `choice`, `noul` (yes/no), or `score` questions. Each batch is one forward pass; large requests may need several batches. **Try it:** [Screen click](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen) · [Space](https://huggingface.co/spaces/thegovind/blink) · [Computer use](https://thegovind.github.io/blink/computer-use/) · [API](https://thegovind.github.io/blink/api/) · [Docs](https://thegovind.github.io/blink/) · [GitHub](https://github.com/thegovind/blink) · [On a Mac (MLX)](#run-it-on-a-mac-mlx) · [blink-4b](https://huggingface.co/thegovind/blink-4b) · [blink-27b](https://huggingface.co/thegovind/blink-27b) ## At a glance | Attribute | Detail | |---|---| | Base model | [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) | | Weights size | 18.8 GB bf16 (8.95B text parameters, 9.41B total) | | Revision | v1.4 (code revision; weights identical to v1.0) | | License | Weights: non-commercial research and evaluation only ([LICENSE.md](LICENSE.md)); code: Apache-2.0. | ## Computer use Screen click uses this model to score numbered targets in a screenshot as a typed choice. ### How it reads a screenshot MiMo's own vision encoder supplies image tokens; blink scores the offered clicks in a forward pass without generating text.
From pixels to probabilities - A screenshot with numbered boxes goes through blink-mimo-9b's own vision encoder from [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B). It has 27 encoder blocks, hidden 1152, and 333 vision tensors unchanged by training. - The encoder turns the screenshot into image tokens. - The language model reads the image tokens plus the typed `choice` question and offered targets. - blink reads next-token logits for the offered options and applies an FP32 softmax. No text is generated. - Enable self-hosted screenshot input with `--vision`.
## Computer use in ten apps blink-mimo-9b finished 133/230 tasks (58%) zero-shot. It's the model behind [Screen click](https://huggingface.co/spaces/thegovind/blink?tab=computer-use&shot=catalog), which runs live in the Space. It uses its own vision encoder with `--vision`, not a borrowed tower. In the "Teach it your app" v1 experiment, 150 labelled screens lifted success on practised apps from 96/160 to 132/160 (60% to 83%). But it paused on only 35/76 truly risky clicks there, down from 62/62 before training. Apps it never saw went from 35/60 to 49/60. Screenshot training is still an experiment, not in the public trainer. These are original apps built for this showcase, not a public benchmark. [Watch the runs](https://huggingface.co/spaces/thegovind/blink?tab=computer-use) · [Videos and traces](https://huggingface.co/datasets/thegovind/blink-cua) · [Code](https://github.com/thegovind/blink/tree/main/examples/cua). [![blink-mimo-9b working on the drawn desktop: a canvas desktop with a dock of app icons for Files, Photos, Mail, Music, Notes and Settings.](assets/blink-mimo-9b-cua.png)](https://huggingface.co/spaces/thegovind/blink?tab=computer-use&scenario=desktop&model=blink-mimo-9b) | App | Tasks done | |---|---:| | Overall | 133/230 (58%) | | Settings | 30/30 | | Desktop (canvas) | 15/20 | | Files | 22/30 | ## Screenshots (opt-in, self-hosted) Self-hosted image input is off by default: start `serve.py` with `--vision` to use MiMo's unchanged tower (image mode needs `torchvision==0.28.0`). TypeSafe's hosted Jev is text-only; screenshots are a blink self-hosted extension, not a hosted Jev feature. Send a `data:image/png;base64,...` URI inside a string in `state`, or data URIs in a top-level `images` list, with a `choice` question whose criteria name the clickable targets. JPEG and WebP data URIs also work. Image URLs are never fetched. These are **development readouts, not a benchmark**. With the default `first` image layout and five offered targets, target accuracy was: | Development set | K5 target accuracy | |---|---:| | ScreenSpot-v2 | 94.10% | | GUIOdyssey labelled subset | 78.95% | Target accuracy is not completed-task success. The default layout was picked on previously seen development data. Try [catalog](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen&shot=catalog), [directory](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen&shot=directory), [lookup](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen&shot=lookup), or [done](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen&shot=done); [Next click (text)](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=nextclick) shows the page-elements-as-text path. For text-only browser agents using page elements as text, [blink-27b](https://huggingface.co/thegovind/blink-27b) did best in local runs; see [browser-agent setup](https://thegovind.github.io/blink/api/#browser-agents). These local runs are not a benchmark. ## Quickstart This example uses text. For screenshots, use the opt-in server path above. ```python # pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub import os, sys from huggingface_hub import hf_hub_download os.environ["BLINK_MODEL"] = "thegovind/blink-mimo-9b" os.environ["BLINK_REVISION"] = "v1.4" sys.path.insert(0, os.path.dirname(hf_hub_download("thegovind/blink-mimo-9b", "blink.py", revision="v1.4"))) import blink out = blink.decide( "Order #4411 arrived with a cracked screen. I want my money back, not another one.", { "intent": { "type": "choice", "instructions": "What does the customer want?", "criteria": {"refund": "Money back", "replacement": "A new unit", "info": "Information only"}, }, "urgent": {"type": "noul", "instructions": "Does this need a reply today?"}, "anger": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]}, }, ) print(out["answers"]["intent"]["probabilities"]) ``` ## Run it as a server `serve.py` implements TypeSafe's text decision API at `POST /v1/systemone` and `GET /v1/models`. Point server-side Python or JavaScript SDKs at it with `TYPESAFE_BASE_URL`. It processes requests one at a time by default; use `--batch-window-ms 5` for cross-request batching. `GET /healthz` reports readiness. The command below starts text-only; pass `--vision` for screenshots. ```sh pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub hf download thegovind/blink-mimo-9b --revision v1.4 --local-dir blink-mimo-9b python blink-mimo-9b/serve.py --model ./blink-mimo-9b --port 8000 # TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any ``` Or use Docker from the downloaded folder: ```sh cd blink-mimo-9b docker build -t blink-mimo-9b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-mimo-9b ``` Earlier vLLM checks passed, but the final read matched earlier answers and still failed hard ECE (0.122 vs 0.121). No MiMo opt-in server ships; use `serve.py`. ## Run it on a Mac (MLX) [`blink_mlx.py`](https://github.com/thegovind/blink/tree/main/examples/mlx) keeps this repo's `blink.py` for prompts, option labels and answers, and runs the forward pass with MLX on Apple silicon. Text and JSON state only, bf16 weights only; it is not a server. It loads only the text side of this checkpoint. ```sh pip install "mlx-lm>=0.31.3" curl -O https://raw.githubusercontent.com/thegovind/blink/main/examples/mlx/blink_mlx.py python blink_mlx.py --model thegovind/blink-mimo-9b ``` `--check` compares your Mac's answers with the Space's saved runs. The [MLX guide](https://github.com/thegovind/blink/tree/main/examples/mlx) covers memory and limits. ## Results | Local development readout | Result | |---|---:| | Decision Index 0.2 balanced skill | 43.36 | | JevBench public hard items | 77/111; hard ECE 0.136 | No official JevBench score for blink has been published. The JevBench numbers are public-item development proxies, not official scores or claims of rank or parity. Decision Index is a descriptive local run of the official kit, not a leaderboard submission; public training exposure affects its scores.
Model details: architecture, training, data ### Architecture and readout ![blink-mimo-9b one-pass readout diagram: a state and typed choice, noul and score questions are rendered as evidence, criterion and lettered options; questions are batched and each batch is one forward pass, giving next-token logits at the answer position, and an FP32 softmax over the offered letters gives one probability per option. No text is generated.](assets/blink-mimo-9b-readout.png) ![blink-mimo-9b network diagram: 32 decoder layers repeating 3 Gated DeltaNet layers then one full-attention layer (24 and 8 in total, hidden 4096), with full attention at 0-based layers 3, 7, … 31 as in the tensor names; LoRA rank 16 on every attention, Gated DeltaNet and MLP projection (43.3M parameters, merged after training); token embeddings, norms and lm_head frozen, with lm_head untied, a separate matrix; the 27-block vision tower kept byte-for-byte; and the answer read from the offered option-letter rows of lm_head.](assets/blink-mimo-9b-network.png) MiMo's text side has 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 4096, untied embeddings, and 8.95B text parameters. The tower has 27 encoder blocks, hidden 1152; all 333 vision tensors unchanged. Fine-tuning merged 43.3M LoRA parameters at rank 16, alpha 32; 179 other language-model tensors stayed unchanged. Targets: `q_proj`, `k_proj`, `v_proj`, `o_proj`; `in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `out_proj`; `gate_proj`, `up_proj`, `down_proj`. Readout uses FP32 softmax on next-token logits for offered labels only. Probabilities are conditional on the options, not certified chances of success. ### Training ![blink-mimo-9b post-training diagram: one run (123,195 rows) broken down by data category feed supervised fine-tuning with cross-entropy over the offered option letters; a single pre-registered epoch is merged and grafted back into the full vision-language checkpoint, vision tower unchanged, as release v1.0. Decision Index 0.1 full suite 56.53.](assets/blink-mimo-9b-post-training.png) One supervised epoch: 123,195 question rows, lr 5e-5, 615 steps. Cross-entropy trained on target distributions; no RL or preference optimization. | Stage | Question rows | Mix | |---|---:|---| | MiMo | 123,195 | 61,394 public-source · 23,894 program-generated reasoning · 12,000 decision worlds · 7,860 teacher-written rows · 7,000 judge-style · 6,000 chess move choices · 5,047 exact-probability worlds | ### Data sources and licences | Source | Licence | |---|---| | MMLU auxiliary train, CommonsenseQA, GSM8K | MIT | | AQuA-RAT, Amazon ESCI | Apache-2.0 | | searchless_chess | data CC BY 4.0 (Lichess-derived portions CC0); code Apache-2.0 | | MedMCQA | Apache-2.0 (dataset card) | | SuperGPQA | ODC-BY | | WANLI, ContractNLI, BANKING77 | CC BY 4.0 | | ARC | CC BY-SA 4.0 | | BoolQ, Dolly-15k | CC BY-SA 3.0 | | ANLI | CC BY-NC 4.0 | | SciQ | CC BY-NC 3.0 | | iSarcasmEval | MIT (upstream repository licence) | | VAST, Humicroedit, OpenBookQA | None stated by source | | Code-generated worlds and teacher-written documents (Qwen3.8-27B) | See LICENSE.md | Source-repository licences do not settle rights in underlying texts. ### Evaluation and limits The DI-S selection sample scored 55.3; the archived Decision Index 0.1 local run scored 56.53. No MMLU-Pro or GPQA items were direct training sources, though SuperGPQA rows matched some added-request text. Public train splits overlap Decision Index, and semantic or pretraining overlap cannot be ruled out. The v1.2 `blink.py` and `serve.py` used its text side only. The screenshot readouts above are development measurements, not final held-out image evaluations. Use `--image-layout first|inline` to place image tokens (`first` is the default), or `--model-name blink-mimo-9b` for a renamed folder. Neither flag switches on images. Text limits are 255 options per choice and 2–10 score levels; invalid images or over-limit requests return 422. English-centric; does not chat or explain answers, and long policies and date/number reasoning remain weak spots.
## License Code: Apache-2.0. Weights: non-commercial research and evaluation only; see each model card's license. See [LICENSE.md](LICENSE.md) for the weight terms. The MiMo model card declares MIT without a separate upstream licence file ([LICENSE-MiMo.md](LICENSE-MiMo.md)); its Qwen base is Apache-2.0 (`LICENSE-Qwen`).