Text Classification
Transformers
Safetensors
English
qwen3_5_text
text-generation
decision-model
typed-decisions
one-pass
option-probabilities
Instructions to use thegovind/blink-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use thegovind/blink-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="thegovind/blink-4b")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("thegovind/blink-4b") model = AutoModelForCausalLM.from_pretrained("thegovind/blink-4b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 10,915 Bytes
656d44d a2e49b3 656d44d 5fbc1e9 846ab75 a2e49b3 656d44d a2e49b3 5fbc1e9 a2e49b3 f7f0e34 2751bf4 5fbc1e9 656d44d f7f0e34 656d44d 5fbc1e9 846ab75 656d44d f7f0e34 656d44d 424f900 656d44d 5fbc1e9 656d44d f7f0e34 400f1f6 5fbc1e9 400f1f6 5fbc1e9 400f1f6 5fbc1e9 400f1f6 5fbc1e9 400f1f6 5fbc1e9 400f1f6 5fbc1e9 400f1f6 5fbc1e9 400f1f6 5fbc1e9 a2e49b3 5fbc1e9 656d44d a2e49b3 656d44d 5fbc1e9 a2e49b3 5fbc1e9 a2e49b3 5fbc1e9 a2e49b3 5fbc1e9 656d44d 5fbc1e9 656d44d a2e49b3 656d44d 5fbc1e9 656d44d 5fbc1e9 656d44d 5fbc1e9 656d44d 5fbc1e9 656d44d a2e49b3 656d44d 400f1f6 5fbc1e9 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 | ---
license: other
license_name: blink-research
license_link: LICENSE.md
base_model: Qwen/Qwen3.5-4B
library_name: transformers
pipeline_tag: text-classification
inference: false
tags:
- decision-model
- typed-decisions
- one-pass
- option-probabilities
language:
- en
---
# blink-4b

Small, fast decisions for routing and checks at volume. Send text or JSON `state` with `choice`, `noul` (yes/no), or `score` questions. Get a probability for every offered answer, not generated text. Each batch takes one forward pass; large requests can use several batches.
Use `choice` to route a request, `noul` for a yes/no check, or `score` for an ordered rating. The same call can ask several questions about a single state.
**Try it:** [Space](https://huggingface.co/spaces/thegovind/blink) · [Screen click](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen) · [Computer use](https://thegovind.github.io/blink/computer-use/) · [API](https://thegovind.github.io/blink/api/) · [Docs](https://thegovind.github.io/blink/) · [GitHub](https://github.com/thegovind/blink) · [blink-mimo-9b](https://huggingface.co/thegovind/blink-mimo-9b) · [blink-27b](https://huggingface.co/thegovind/blink-27b)
## At a glance
| Attribute | Detail |
|---|---|
| Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B), text weights only |
| Weights size | 8.4 GB (bf16, 4,205,751,296 parameters) |
| Revision | v1.4 (code revision; weights identical to v1.0) |
| License | Weights: non-commercial research and evaluation only ([LICENSE.md](LICENSE.md)); code: Apache-2.0. Base-model notice: Apache-2.0 (`LICENSE-Qwen`). |
## Quickstart
```python
# pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
import os, sys
from huggingface_hub import hf_hub_download
os.environ["BLINK_MODEL"] = "thegovind/blink-4b"
os.environ["BLINK_REVISION"] = "v1.4"
sys.path.insert(0, os.path.dirname(hf_hub_download("thegovind/blink-4b", "blink.py", revision="v1.4")))
import blink
out = blink.decide(
"Order #4411 arrived with a cracked screen. I want my money back, not another one.",
{
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": {"refund": "Money back", "replacement": "A new unit", "info": "Information only"},
},
"urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
"anger": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]},
},
)
print(out["answers"]["intent"]["probabilities"])
```
## Run it as a server
`serve.py` serves `POST /v1/systemone`, `GET /v1/models`, and `GET /healthz`. Point TypeSafe's server-side Python or JavaScript SDKs at it with `TYPESAFE_BASE_URL`; text decisions use the same request and response fields as hosted Jev. Requests run one at a time by default; `--batch-window-ms 5` enables cross-request batching. The [API reference](https://thegovind.github.io/blink/api/) covers limits and errors.
```sh
pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
hf download thegovind/blink-4b --revision v1.4 --local-dir blink-4b
python blink-4b/serve.py --model ./blink-4b --port 8000
# TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any
```
Or use Docker from the downloaded folder:
```sh
cd blink-4b
docker build -t blink-4b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-4b
```
## Higher throughput (opt-in)
`serve_vllm.py` (added in `v1.4`) is an opt-in, text-only server with higher throughput. `serve.py` stays the default.
Earlier paired loopback measurements (one self-hosted replica, 231 public items): c1 p50/p95 was 52/127 ms on vLLM vs 64/175 ms on `serve.py` (one request at a time); c16 was 34.7 vs 12.2 completed questions/s (16 concurrent requests).
Earlier peak memory was 68.0 GiB including cache and cold start 180 s; neither was retimed, nor were the other quality sets or c16 rerun. A release build re-passed the 231-item JevBench c1 quality check (80/111 hard, 199/231 total, hard ECE 0.069). Later changes touched only request checks, error handling and startup cleanup, not scoring.
Tested versions:
```sh
python -m pip install "torch==2.13.0" "transformers==5.17.0" \
"vllm==0.30.0" "compressed-tensors==0.17.0" \
"accelerate>=1.1.0" safetensors huggingface_hub
python -m pip install "flash-linear-attention==0.5.2"
```
```sh
hf download thegovind/blink-4b --revision v1.4 --local-dir blink-4b
cd blink-4b
python serve_vllm.py --model . --port 8000 --quantization auto --max-concurrency 32 --max-num-seqs 32 --max-num-batched-tokens 8192 --gpu-memory-utilization 0.85
```
[Setup, quality limits and caveats](https://huggingface.co/thegovind/blink-4b/blob/v1.4/VLLM.md). No official JevBench score for blink has been published.
## Screenshots (opt-in, self-hosted)
Image input is off by default. Start `serve.py` with `--vision-tower Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` to attach the matching tower. This is a self-hosted blink extension. TypeSafe's hosted Jev is text-only.
With `--vision-tower`, self-hosted blink borrows the pinned Qwen base model's vision encoder while its checkpoint stays text-only; [blink-mimo-9b](https://huggingface.co/thegovind/blink-mimo-9b) uses its own encoder and runs [Screen click](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen).
Put a `data:image/png;base64,...` URI (JPEG and WebP data URIs work too) inside a `state` string, or pass data URIs in a top-level `images` list. Image URLs are never fetched. If loading only `blink.py` via `hf_hub_download`, also download `graft_keys.py` from the same repo and revision beside it.
See [Computer use](https://thegovind.github.io/blink/computer-use/) to self-host this model with images.
## Results
| Local development readout | Result |
|---|---:|
| JevBench public-items proxy | 76.5; 80/111 hard; hard ECE 0.067 |
| Decision Index 0.2 balanced skill | 37.85 |
No official JevBench score for blink has been published. The JevBench numbers are public-item development proxies, not an official score, rank, or parity claim. Decision Index is a descriptive local run of the official kit, not a leaderboard submission; training overlap affects its scores.
<details><summary>Model details: architecture, training, data</summary>
### Architecture and readout

Qwen3.5-4B text backbone: 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 2560, and 4,205,751,296 shipped parameters. The training updated 32.5M LoRA parameters at rank 16, alpha 32: `q_proj`, `k_proj`, `v_proj`, `o_proj`; `in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `out_proj`; and `gate_proj`, `up_proj`, `down_proj`. The vision encoder and MTP head were removed: 0 vision tensors, 0 MTP tensors.
Readout softmaxes FP32 next-token logits over the offered option labels. These are option-conditional probabilities, not certified chances of being right.
### Training

Supervised fine-tuning on target distributions; no RL or preference optimization. T3 used lr 3e-5 for 96 steps; T4 used lr 4e-5 for 472 steps. The released weights average T3, T4 step 300, and T4 final.
| Stage | Question rows | Mix |
|---|---:|---|
| T3 | 23,156 | 11,352 decision worlds · 3,741 teacher-written rows · 2,579 exact-probability worlds · 2,221 public-source · 1,763 base-model anchors · 1,500 program-generated reasoning |
| T4 | 42,360 | 12,000 decision worlds · 7,860 teacher-written rows · 7,000 judge-style · 6,220 exact-probability worlds · 3,500 base-model anchors · 3,000 program-generated reasoning · 2,780 public-source |
### Data sources and licences
| Source | Licence |
|---|---|
| MMLU auxiliary train, MMLU-Pro, CommonsenseQA, GSM8K | MIT |
| AQuA-RAT, Amazon ESCI | Apache-2.0 |
| MedMCQA | Apache-2.0 (dataset card) |
| SuperGPQA | ODC-BY |
| WANLI, ContractNLI, BANKING77, GPQA | CC BY 4.0 |
| ARC | CC BY-SA 4.0 |
| BoolQ, Dolly-15k | CC BY-SA 3.0 |
| ANLI | CC BY-NC 4.0 |
| SciQ | CC BY-NC 3.0 |
| iSarcasmEval | MIT (upstream repository licence) |
| VAST, Humicroedit, OpenBookQA | None stated by source |
| Code-generated worlds and teacher-written documents (Qwen3.8-27B) | See LICENSE.md |
Source-repository licences do not settle rights in every underlying text.
### Evaluation and limits
The archived Decision Index 0.1 local run scored 52.12; editions are not directly comparable. Public JevBench items were used for development selection, not training. Training included 281 MMLU-Pro test-partition questions, contaminating that Decision Index 0.2 component; semantic and pretraining overlap cannot be ruled out. English-centric, no chatting or explanations; text in `state` can sway an answer, especially on long policies and date arithmetic.
For `blink.py` and default `serve.py`: 255 options per choice, 2–10 score levels, 131,072 input tokens per question, and 512 questions per request; over-limit requests return 422 without truncation. For image placement use `--image-layout first|inline` (`first` is the default); for a renamed folder use `--model-name blink-4b`. Neither flag enables images on its own.
</details>
## License
Code: Apache-2.0. Weights: non-commercial research and evaluation only; see each model card's license.
See [LICENSE.md](LICENSE.md) for weight terms; the Qwen base model is Apache-2.0 (`LICENSE-Qwen`).
|