How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "thegovind/blink-mimo-9b"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "thegovind/blink-mimo-9b",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Use Docker
docker model run hf.co/thegovind/blink-mimo-9b
Quick Links

blink-mimo-9b

blink-mimo-9b Screen click diagram: the “Open the result” preset, a screenshot of a book catalog page with 3 numbered boxes: box 1 on the select “Book category: Maps”, box 2 on the button “Search catalog” and box 3 on the link “Open Atlas Field Guide”. The task is “Search the Maps category, then open Atlas Field Guide from the results.” The screenshot is resized to multiples of 32 px (1440 × 864) and read by blink-mimo-9b's own vision encoder from XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B (27 blocks, hidden 1152 → 4096; its 333 tensors unchanged by blink training), which turns it into 1,215 image tokens (45 × 27, one per 32 × 32 px). Three typed questions are read with the image tokens, one forward pass each, with no text generated: which box to act on next (a choice over boxes 1 to 3), is the task already done (yes or no), and would the next click be risky (yes or no). In this saved run, as the Space shows it (saved run · blink-mimo-9b), box 3 has the top probability, 62%, then box 2 33% and box 1 5%; done is 19% yes and risky 2% yes, so the verdict is “Click 3”. Demo only; it never clicks.

Pick the next click from a screenshot. blink-mimo-9b reads it with its own vision tower and scores the offered targets as a typed choice, returning probabilities instead of generated text. It also takes text or JSON state with choice, noul (yes/no), or score questions. Each batch is one forward pass; large requests may need several batches.

Try it: Screen click · Space · Computer use · API · Docs · GitHub · blink-4b · blink-27b

At a glance

Attribute Detail
Base model XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Weights size 18.8 GB bf16 (8.95B text parameters, 9.41B total)
Revision v1.4 (code revision; weights identical to v1.0)
License Weights: non-commercial research and evaluation only (LICENSE.md); code: Apache-2.0.

Computer use

Screen click uses this model to score numbered targets in a screenshot as a typed choice.

How it reads a screenshot

MiMo's own vision encoder supplies image tokens; blink scores the offered clicks in a forward pass without generating text.

From pixels to probabilities
  • A screenshot with numbered boxes goes through blink-mimo-9b's own vision encoder from XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B. It has 27 encoder blocks, hidden 1152, and 333 vision tensors unchanged by training.
  • The encoder turns the screenshot into image tokens.
  • The language model reads the image tokens plus the typed choice question and offered targets.
  • blink reads next-token logits for the offered options and applies an FP32 softmax. No text is generated.
  • Enable self-hosted screenshot input with --vision.

Computer use in ten apps

blink-mimo-9b finished 133/230 tasks (58%) zero-shot. It's the model behind Screen click, which runs live in the Space. It uses its own vision encoder with --vision, not a borrowed tower. In the "Teach it your app" v1 experiment, 150 labelled screens lifted success on practised apps from 96/160 to 132/160 (60% to 83%). But it paused on only 35/76 truly risky clicks there, down from 62/62 before training. Apps it never saw went from 35/60 to 49/60. Screenshot training is still an experiment, not in the public trainer. These are original apps built for this showcase, not a public benchmark. Watch the runs · Videos and traces · Code.

blink-mimo-9b working on the drawn desktop: a canvas desktop with a dock of app icons for Files, Photos, Mail, Music, Notes and Settings.

App Tasks done
Overall 133/230 (58%)
Settings 30/30
Desktop (canvas) 15/20
Files 22/30

Screenshots (opt-in, self-hosted)

Self-hosted image input is off by default: start serve.py with --vision to use MiMo's unchanged tower (image mode needs torchvision==0.28.0). TypeSafe's hosted Jev is text-only; screenshots are a blink self-hosted extension, not a hosted Jev feature.

Send a data:image/png;base64,... URI inside a string in state, or data URIs in a top-level images list, with a choice question whose criteria name the clickable targets. JPEG and WebP data URIs also work. Image URLs are never fetched.

These are development readouts, not a benchmark. With the default first image layout and five offered targets, target accuracy was:

Development set K5 target accuracy
ScreenSpot-v2 94.10%
GUIOdyssey labelled subset 78.95%

Target accuracy is not completed-task success. The default layout was picked on previously seen development data. Try catalog, directory, lookup, or done; Next click (text) shows the page-elements-as-text path.

For text-only browser agents using page elements as text, blink-27b did best in local runs; see browser-agent setup. These local runs are not a benchmark.

Quickstart

This example uses text. For screenshots, use the opt-in server path above.

# pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
import os, sys
from huggingface_hub import hf_hub_download

os.environ["BLINK_MODEL"] = "thegovind/blink-mimo-9b"
os.environ["BLINK_REVISION"] = "v1.4"
sys.path.insert(0, os.path.dirname(hf_hub_download("thegovind/blink-mimo-9b", "blink.py", revision="v1.4")))
import blink

out = blink.decide(
    "Order #4411 arrived with a cracked screen. I want my money back, not another one.",
    {
        "intent": {
            "type": "choice",
            "instructions": "What does the customer want?",
            "criteria": {"refund": "Money back", "replacement": "A new unit", "info": "Information only"},
        },
        "urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
        "anger": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]},
    },
)
print(out["answers"]["intent"]["probabilities"])

Run it as a server

serve.py implements TypeSafe's text decision API at POST /v1/systemone and GET /v1/models. Point server-side Python or JavaScript SDKs at it with TYPESAFE_BASE_URL. It processes requests one at a time by default; use --batch-window-ms 5 for cross-request batching. GET /healthz reports readiness. The command below starts text-only; pass --vision for screenshots.

pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
hf download thegovind/blink-mimo-9b --revision v1.4 --local-dir blink-mimo-9b
python blink-mimo-9b/serve.py --model ./blink-mimo-9b --port 8000
# TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any

Or use Docker from the downloaded folder:

cd blink-mimo-9b
docker build -t blink-mimo-9b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-mimo-9b

Earlier vLLM checks passed, but the final read matched earlier answers and still failed hard ECE (0.122 vs 0.121). No MiMo opt-in server ships; use serve.py.

Results

Local development readout Result
Decision Index 0.2 balanced skill 43.36
JevBench public hard items 77/111; hard ECE 0.136

No official JevBench score for blink has been published. The JevBench numbers are public-item development proxies, not official scores or claims of rank or parity. Decision Index is a descriptive local run of the official kit, not a leaderboard submission; public training exposure affects its scores.

Model details: architecture, training, data

Architecture and readout

blink-mimo-9b one-pass readout diagram: a state and typed choice, noul and score questions are rendered as evidence, criterion and lettered options; questions are batched and each batch is one forward pass, giving next-token logits at the answer position, and an FP32 softmax over the offered letters gives one probability per option. No text is generated.

blink-mimo-9b network diagram: 32 decoder layers repeating 3 Gated DeltaNet layers then one full-attention layer (24 and 8 in total, hidden 4096), with full attention at 0-based layers 3, 7, … 31 as in the tensor names; LoRA rank 16 on every attention, Gated DeltaNet and MLP projection (43.3M parameters, merged after training); token embeddings, norms and lm_head frozen, with lm_head untied, a separate matrix; the 27-block vision tower kept byte-for-byte; and the answer read from the offered option-letter rows of lm_head.

MiMo's text side has 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 4096, untied embeddings, and 8.95B text parameters. The tower has 27 encoder blocks, hidden 1152; all 333 vision tensors unchanged. Fine-tuning merged 43.3M LoRA parameters at rank 16, alpha 32; 179 other language-model tensors stayed unchanged. Targets: q_proj, k_proj, v_proj, o_proj; in_proj_qkv, in_proj_z, in_proj_a, in_proj_b, out_proj; gate_proj, up_proj, down_proj.

Readout uses FP32 softmax on next-token logits for offered labels only. Probabilities are conditional on the options, not certified chances of success.

Training

blink-mimo-9b post-training diagram: one run (123,195 rows) broken down by data category feed supervised fine-tuning with cross-entropy over the offered option letters; a single pre-registered epoch is merged and grafted back into the full vision-language checkpoint, vision tower unchanged, as release v1.0. Decision Index 0.1 full suite 56.53.

One supervised epoch: 123,195 question rows, lr 5e-5, 615 steps. Cross-entropy trained on target distributions; no RL or preference optimization.

Stage Question rows Mix
MiMo 123,195 61,394 public-source · 23,894 program-generated reasoning · 12,000 decision worlds · 7,860 teacher-written rows · 7,000 judge-style · 6,000 chess move choices · 5,047 exact-probability worlds

Data sources and licences

Source Licence
MMLU auxiliary train, CommonsenseQA, GSM8K MIT
AQuA-RAT, Amazon ESCI Apache-2.0
searchless_chess data CC BY 4.0 (Lichess-derived portions CC0); code Apache-2.0
MedMCQA Apache-2.0 (dataset card)
SuperGPQA ODC-BY
WANLI, ContractNLI, BANKING77 CC BY 4.0
ARC CC BY-SA 4.0
BoolQ, Dolly-15k CC BY-SA 3.0
ANLI CC BY-NC 4.0
SciQ CC BY-NC 3.0
iSarcasmEval MIT (upstream repository licence)
VAST, Humicroedit, OpenBookQA None stated by source
Code-generated worlds and teacher-written documents (Qwen3.8-27B) See LICENSE.md

Source-repository licences do not settle rights in underlying texts.

Evaluation and limits

The DI-S selection sample scored 55.3; the archived Decision Index 0.1 local run scored 56.53. No MMLU-Pro or GPQA items were direct training sources, though SuperGPQA rows matched some added-request text. Public train splits overlap Decision Index, and semantic or pretraining overlap cannot be ruled out. The v1.2 blink.py and serve.py used its text side only. The screenshot readouts above are development measurements, not final held-out image evaluations.

Use --image-layout first|inline to place image tokens (first is the default), or --model-name blink-mimo-9b for a renamed folder. Neither flag switches on images. Text limits are 255 options per choice and 2–10 score levels; invalid images or over-limit requests return 422. English-centric; does not chat or explain answers, and long policies and date/number reasoning remain weak spots.

License

Code: Apache-2.0. Weights: non-commercial research and evaluation only; see each model card's license.

See LICENSE.md for the weight terms. The MiMo model card declares MIT without a separate upstream licence file (LICENSE-MiMo.md); its Qwen base is Apache-2.0 (LICENSE-Qwen).

Downloads last month
140
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thegovind/blink-mimo-9b

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(20)
this model
Quantizations
1 model

Space using thegovind/blink-mimo-9b 1