Instructions to use thegovind/blink-mimo-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use thegovind/blink-mimo-9b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="thegovind/blink-mimo-9b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("thegovind/blink-mimo-9b") model = AutoModelForMultimodalLM.from_pretrained("thegovind/blink-mimo-9b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use thegovind/blink-mimo-9b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "thegovind/blink-mimo-9b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thegovind/blink-mimo-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/thegovind/blink-mimo-9b
- SGLang
How to use thegovind/blink-mimo-9b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "thegovind/blink-mimo-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thegovind/blink-mimo-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "thegovind/blink-mimo-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thegovind/blink-mimo-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use thegovind/blink-mimo-9b with Docker Model Runner:
docker model run hf.co/thegovind/blink-mimo-9b
blink-mimo-9b
Pick the next click from a screenshot. blink-mimo-9b reads it with its own vision tower and scores the offered targets as a typed choice, returning probabilities instead of generated text. It also takes text or JSON state with choice, noul (yes/no), or score questions. Each batch is one forward pass; large requests may need several batches.
Try it: Screen click · Space · Computer use · API · Docs · GitHub · blink-4b · blink-27b
At a glance
| Attribute | Detail |
|---|---|
| Base model | XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B |
| Weights size | 18.8 GB bf16 (8.95B text parameters, 9.41B total) |
| Revision | v1.4 (code revision; weights identical to v1.0) |
| License | Weights: non-commercial research and evaluation only (LICENSE.md); code: Apache-2.0. |
Computer use
Screen click uses this model to score numbered targets in a screenshot as a typed choice.
How it reads a screenshot
MiMo's own vision encoder supplies image tokens; blink scores the offered clicks in a forward pass without generating text.
From pixels to probabilities
- A screenshot with numbered boxes goes through blink-mimo-9b's own vision encoder from XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B. It has 27 encoder blocks, hidden 1152, and 333 vision tensors unchanged by training.
- The encoder turns the screenshot into image tokens.
- The language model reads the image tokens plus the typed
choicequestion and offered targets. - blink reads next-token logits for the offered options and applies an FP32 softmax. No text is generated.
- Enable self-hosted screenshot input with
--vision.
Computer use in ten apps
blink-mimo-9b finished 133/230 tasks (58%) zero-shot. It's the model behind Screen click, which runs live in the Space. It uses its own vision encoder with --vision, not a borrowed tower. In the "Teach it your app" v1 experiment, 150 labelled screens lifted success on practised apps from 96/160 to 132/160 (60% to 83%). But it paused on only 35/76 truly risky clicks there, down from 62/62 before training. Apps it never saw went from 35/60 to 49/60. Screenshot training is still an experiment, not in the public trainer. These are original apps built for this showcase, not a public benchmark. Watch the runs · Videos and traces · Code.
| App | Tasks done |
|---|---|
| Overall | 133/230 (58%) |
| Settings | 30/30 |
| Desktop (canvas) | 15/20 |
| Files | 22/30 |
Screenshots (opt-in, self-hosted)
Self-hosted image input is off by default: start serve.py with --vision to use MiMo's unchanged tower (image mode needs torchvision==0.28.0). TypeSafe's hosted Jev is text-only; screenshots are a blink self-hosted extension, not a hosted Jev feature.
Send a data:image/png;base64,... URI inside a string in state, or data URIs in a top-level images list, with a choice question whose criteria name the clickable targets. JPEG and WebP data URIs also work. Image URLs are never fetched.
These are development readouts, not a benchmark. With the default first image layout and five offered targets, target accuracy was:
| Development set | K5 target accuracy |
|---|---|
| ScreenSpot-v2 | 94.10% |
| GUIOdyssey labelled subset | 78.95% |
Target accuracy is not completed-task success. The default layout was picked on previously seen development data. Try catalog, directory, lookup, or done; Next click (text) shows the page-elements-as-text path.
For text-only browser agents using page elements as text, blink-27b did best in local runs; see browser-agent setup. These local runs are not a benchmark.
Quickstart
This example uses text. For screenshots, use the opt-in server path above.
# pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
import os, sys
from huggingface_hub import hf_hub_download
os.environ["BLINK_MODEL"] = "thegovind/blink-mimo-9b"
os.environ["BLINK_REVISION"] = "v1.4"
sys.path.insert(0, os.path.dirname(hf_hub_download("thegovind/blink-mimo-9b", "blink.py", revision="v1.4")))
import blink
out = blink.decide(
"Order #4411 arrived with a cracked screen. I want my money back, not another one.",
{
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": {"refund": "Money back", "replacement": "A new unit", "info": "Information only"},
},
"urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
"anger": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]},
},
)
print(out["answers"]["intent"]["probabilities"])
Run it as a server
serve.py implements TypeSafe's text decision API at POST /v1/systemone and GET /v1/models. Point server-side Python or JavaScript SDKs at it with TYPESAFE_BASE_URL. It processes requests one at a time by default; use --batch-window-ms 5 for cross-request batching. GET /healthz reports readiness. The command below starts text-only; pass --vision for screenshots.
pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
hf download thegovind/blink-mimo-9b --revision v1.4 --local-dir blink-mimo-9b
python blink-mimo-9b/serve.py --model ./blink-mimo-9b --port 8000
# TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any
Or use Docker from the downloaded folder:
cd blink-mimo-9b
docker build -t blink-mimo-9b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-mimo-9b
Earlier vLLM checks passed, but the final read matched earlier answers and still failed hard ECE (0.122 vs 0.121). No MiMo opt-in server ships; use serve.py.
Results
| Local development readout | Result |
|---|---|
| Decision Index 0.2 balanced skill | 43.36 |
| JevBench public hard items | 77/111; hard ECE 0.136 |
No official JevBench score for blink has been published. The JevBench numbers are public-item development proxies, not official scores or claims of rank or parity. Decision Index is a descriptive local run of the official kit, not a leaderboard submission; public training exposure affects its scores.
Model details: architecture, training, data
Architecture and readout
MiMo's text side has 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 4096, untied embeddings, and 8.95B text parameters. The tower has 27 encoder blocks, hidden 1152; all 333 vision tensors unchanged. Fine-tuning merged 43.3M LoRA parameters at rank 16, alpha 32; 179 other language-model tensors stayed unchanged. Targets: q_proj, k_proj, v_proj, o_proj; in_proj_qkv, in_proj_z, in_proj_a, in_proj_b, out_proj; gate_proj, up_proj, down_proj.
Readout uses FP32 softmax on next-token logits for offered labels only. Probabilities are conditional on the options, not certified chances of success.
Training
One supervised epoch: 123,195 question rows, lr 5e-5, 615 steps. Cross-entropy trained on target distributions; no RL or preference optimization.
| Stage | Question rows | Mix |
|---|---|---|
| MiMo | 123,195 | 61,394 public-source · 23,894 program-generated reasoning · 12,000 decision worlds · 7,860 teacher-written rows · 7,000 judge-style · 6,000 chess move choices · 5,047 exact-probability worlds |
Data sources and licences
| Source | Licence |
|---|---|
| MMLU auxiliary train, CommonsenseQA, GSM8K | MIT |
| AQuA-RAT, Amazon ESCI | Apache-2.0 |
| searchless_chess | data CC BY 4.0 (Lichess-derived portions CC0); code Apache-2.0 |
| MedMCQA | Apache-2.0 (dataset card) |
| SuperGPQA | ODC-BY |
| WANLI, ContractNLI, BANKING77 | CC BY 4.0 |
| ARC | CC BY-SA 4.0 |
| BoolQ, Dolly-15k | CC BY-SA 3.0 |
| ANLI | CC BY-NC 4.0 |
| SciQ | CC BY-NC 3.0 |
| iSarcasmEval | MIT (upstream repository licence) |
| VAST, Humicroedit, OpenBookQA | None stated by source |
| Code-generated worlds and teacher-written documents (Qwen3.8-27B) | See LICENSE.md |
Source-repository licences do not settle rights in underlying texts.
Evaluation and limits
The DI-S selection sample scored 55.3; the archived Decision Index 0.1 local run scored 56.53. No MMLU-Pro or GPQA items were direct training sources, though SuperGPQA rows matched some added-request text. Public train splits overlap Decision Index, and semantic or pretraining overlap cannot be ruled out. The v1.2 blink.py and serve.py used its text side only. The screenshot readouts above are development measurements, not final held-out image evaluations.
Use --image-layout first|inline to place image tokens (first is the default), or --model-name blink-mimo-9b for a renamed folder. Neither flag switches on images. Text limits are 255 options per choice and 2–10 score levels; invalid images or over-limit requests return 422. English-centric; does not chat or explain answers, and long policies and date/number reasoning remain weak spots.
License
Code: Apache-2.0. Weights: non-commercial research and evaluation only; see each model card's license.
See LICENSE.md for the weight terms. The MiMo model card declares MIT without a separate upstream licence file (LICENSE-MiMo.md); its Qwen base is Apache-2.0 (LICENSE-Qwen).
- Downloads last month
- 140
Model tree for thegovind/blink-mimo-9b
Base model
Qwen/Qwen3.5-9B-Base




Install from pip and serve model
# Install vLLM from pip: pip install vllm# Start the vLLM server: vllm serve "thegovind/blink-mimo-9b"# Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thegovind/blink-mimo-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'