blink-4b

blink-4b one-pass readout diagram: a state and typed choice, noul and score questions are rendered as evidence, criterion and lettered options; questions are batched and each batch is one forward pass, giving next-token logits at the answer position, and an FP32 softmax over the offered letters gives one probability per option. No text is generated.

Small, fast decisions for routing and checks at volume. Send text or JSON state with choice, noul (yes/no), or score questions. Get a probability for every offered answer, not generated text. Each batch takes one forward pass; large requests can use several batches.

Use choice to route a request, noul for a yes/no check, or score for an ordered rating. The same call can ask several questions about a single state.

Try it: Space · Screen click · Computer use · API · Docs · GitHub · blink-mimo-9b · blink-27b

At a glance

Attribute Detail
Base model Qwen/Qwen3.5-4B, text weights only
Weights size 8.4 GB (bf16, 4,205,751,296 parameters)
Revision v1.4 (code revision; weights identical to v1.0)
License Weights: non-commercial research and evaluation only (LICENSE.md); code: Apache-2.0. Base-model notice: Apache-2.0 (LICENSE-Qwen).

Quickstart

# pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
import os, sys
from huggingface_hub import hf_hub_download

os.environ["BLINK_MODEL"] = "thegovind/blink-4b"
os.environ["BLINK_REVISION"] = "v1.4"
sys.path.insert(0, os.path.dirname(hf_hub_download("thegovind/blink-4b", "blink.py", revision="v1.4")))
import blink

out = blink.decide(
    "Order #4411 arrived with a cracked screen. I want my money back, not another one.",
    {
        "intent": {
            "type": "choice",
            "instructions": "What does the customer want?",
            "criteria": {"refund": "Money back", "replacement": "A new unit", "info": "Information only"},
        },
        "urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
        "anger": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]},
    },
)
print(out["answers"]["intent"]["probabilities"])

Run it as a server

serve.py serves POST /v1/systemone, GET /v1/models, and GET /healthz. Point TypeSafe's server-side Python or JavaScript SDKs at it with TYPESAFE_BASE_URL; text decisions use the same request and response fields as hosted Jev. Requests run one at a time by default; --batch-window-ms 5 enables cross-request batching. The API reference covers limits and errors.

pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
hf download thegovind/blink-4b --revision v1.4 --local-dir blink-4b
python blink-4b/serve.py --model ./blink-4b --port 8000
# TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any

Or use Docker from the downloaded folder:

cd blink-4b
docker build -t blink-4b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-4b

Higher throughput (opt-in)

serve_vllm.py (added in v1.4) is an opt-in, text-only server with higher throughput. serve.py stays the default.

Earlier paired loopback measurements (one self-hosted replica, 231 public items): c1 p50/p95 was 52/127 ms on vLLM vs 64/175 ms on serve.py (one request at a time); c16 was 34.7 vs 12.2 completed questions/s (16 concurrent requests).

Earlier peak memory was 68.0 GiB including cache and cold start 180 s; neither was retimed, nor were the other quality sets or c16 rerun. A release build re-passed the 231-item JevBench c1 quality check (80/111 hard, 199/231 total, hard ECE 0.069). Later changes touched only request checks, error handling and startup cleanup, not scoring.

Tested versions:

python -m pip install "torch==2.13.0" "transformers==5.17.0" \
  "vllm==0.30.0" "compressed-tensors==0.17.0" \
  "accelerate>=1.1.0" safetensors huggingface_hub
python -m pip install "flash-linear-attention==0.5.2"
hf download thegovind/blink-4b --revision v1.4 --local-dir blink-4b
cd blink-4b
python serve_vllm.py --model . --port 8000 --quantization auto --max-concurrency 32 --max-num-seqs 32 --max-num-batched-tokens 8192 --gpu-memory-utilization 0.85

Setup, quality limits and caveats. No official JevBench score for blink has been published.

Screenshots (opt-in, self-hosted)

Image input is off by default. Start serve.py with --vision-tower Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a to attach the matching tower. This is a self-hosted blink extension. TypeSafe's hosted Jev is text-only.

With --vision-tower, self-hosted blink borrows the pinned Qwen base model's vision encoder while its checkpoint stays text-only; blink-mimo-9b uses its own encoder and runs Screen click.

Put a data:image/png;base64,... URI (JPEG and WebP data URIs work too) inside a state string, or pass data URIs in a top-level images list. Image URLs are never fetched. If loading only blink.py via hf_hub_download, also download graft_keys.py from the same repo and revision beside it.

See Computer use to self-host this model with images.

Results

Local development readout Result
JevBench public-items proxy 76.5; 80/111 hard; hard ECE 0.067
Decision Index 0.2 balanced skill 37.85

No official JevBench score for blink has been published. The JevBench numbers are public-item development proxies, not an official score, rank, or parity claim. Decision Index is a descriptive local run of the official kit, not a leaderboard submission; training overlap affects its scores.

Model details: architecture, training, data

Architecture and readout

blink-4b network diagram: 32 decoder layers repeating 3 Gated DeltaNet layers then one full-attention layer (24 and 8 in total, hidden 2560), with full attention at 0-based layers 3, 7, … 31 as in the tensor names; LoRA rank 16 on every attention, Gated DeltaNet and MLP projection (32.5M parameters, merged after training); token embeddings, norms and lm_head frozen, with lm_head tied to the token embeddings, one matrix; the vision encoder and multi-token-prediction head removed; and the answer read from the offered option-letter rows of lm_head.

Qwen3.5-4B text backbone: 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 2560, and 4,205,751,296 shipped parameters. The training updated 32.5M LoRA parameters at rank 16, alpha 32: q_proj, k_proj, v_proj, o_proj; in_proj_qkv, in_proj_z, in_proj_a, in_proj_b, out_proj; and gate_proj, up_proj, down_proj. The vision encoder and MTP head were removed: 0 vision tensors, 0 MTP tensors.

Readout softmaxes FP32 next-token logits over the offered option labels. These are option-conditional probabilities, not certified chances of being right.

Training

blink-4b post-training diagram: T3 (23,156 rows) and T4 (42,360 rows) broken down by data category feed supervised fine-tuning with cross-entropy over the offered option letters; both runs start from the base, and T3, T4 step 300 and T4 final are merged and averaged into release v1.0. Decision Index 0.1 full suite 52.12.

Supervised fine-tuning on target distributions; no RL or preference optimization. T3 used lr 3e-5 for 96 steps; T4 used lr 4e-5 for 472 steps. The released weights average T3, T4 step 300, and T4 final.

Stage Question rows Mix
T3 23,156 11,352 decision worlds · 3,741 teacher-written rows · 2,579 exact-probability worlds · 2,221 public-source · 1,763 base-model anchors · 1,500 program-generated reasoning
T4 42,360 12,000 decision worlds · 7,860 teacher-written rows · 7,000 judge-style · 6,220 exact-probability worlds · 3,500 base-model anchors · 3,000 program-generated reasoning · 2,780 public-source

Data sources and licences

Source Licence
MMLU auxiliary train, MMLU-Pro, CommonsenseQA, GSM8K MIT
AQuA-RAT, Amazon ESCI Apache-2.0
MedMCQA Apache-2.0 (dataset card)
SuperGPQA ODC-BY
WANLI, ContractNLI, BANKING77, GPQA CC BY 4.0
ARC CC BY-SA 4.0
BoolQ, Dolly-15k CC BY-SA 3.0
ANLI CC BY-NC 4.0
SciQ CC BY-NC 3.0
iSarcasmEval MIT (upstream repository licence)
VAST, Humicroedit, OpenBookQA None stated by source
Code-generated worlds and teacher-written documents (Qwen3.8-27B) See LICENSE.md

Source-repository licences do not settle rights in every underlying text.

Evaluation and limits

The archived Decision Index 0.1 local run scored 52.12; editions are not directly comparable. Public JevBench items were used for development selection, not training. Training included 281 MMLU-Pro test-partition questions, contaminating that Decision Index 0.2 component; semantic and pretraining overlap cannot be ruled out. English-centric, no chatting or explanations; text in state can sway an answer, especially on long policies and date arithmetic.

For blink.py and default serve.py: 255 options per choice, 2–10 score levels, 131,072 input tokens per question, and 512 questions per request; over-limit requests return 422 without truncation. For image placement use --image-layout first|inline (first is the default); for a renamed folder use --model-name blink-4b. Neither flag enables images on its own.

License

Code: Apache-2.0. Weights: non-commercial research and evaluation only; see each model card's license.

See LICENSE.md for weight terms; the Qwen base model is Apache-2.0 (LICENSE-Qwen).

Downloads last month
234
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thegovind/blink-4b

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(790)
this model

Space using thegovind/blink-4b 1