thegovind commited on
Commit
5fbc1e9
·
verified ·
1 Parent(s): f8d6c2f

Card: crisper card, computer use first

Browse files
Files changed (1) hide show
  1. README.md +32 -193
README.md CHANGED
@@ -17,135 +17,24 @@ language:
17
 
18
  # blink-4b
19
 
20
- Send a text or JSON `state` and typed questions: `choice` picks from up to 255 options, `noul` is yes/no, and `score` takes 2–10 ordered levels. Each question gets probabilities over its offered options from one forward pass, with no generated text. Long or large multi-question requests may use several batches.
21
-
22
  ![blink-4b one-pass readout diagram: a state and typed choice, noul and score questions are rendered as evidence, criterion and lettered options; questions are batched and each batch is one forward pass, giving next-token logits at the answer position, and an FP32 softmax over the offered letters gives one probability per option. No text is generated.](assets/blink-4b-readout.png)
23
 
24
- **Try it:** [Space demo](https://huggingface.co/spaces/thegovind/blink) · [blink-4b](https://huggingface.co/thegovind/blink-4b) · [blink-27b](https://huggingface.co/thegovind/blink-27b) · [blink-mimo-9b](https://huggingface.co/thegovind/blink-mimo-9b) · [Source code](https://github.com/thegovind/blink) · [Docs](https://thegovind.github.io/blink/) · [API](https://thegovind.github.io/blink/api/)
 
 
 
 
25
 
26
  ## At a glance
27
 
28
  | Attribute | Detail |
29
  |---|---|
30
- | Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (text model only; vision encoder and MTP head removed) |
31
  | Weights size | 8.4 GB (bf16, 4,205,751,296 parameters) |
32
  | Revision | v1.4 (code revision; weights identical to v1.0) |
33
  | License | Weights: non-commercial research and evaluation only ([LICENSE.md](LICENSE.md)); code: Apache-2.0. Base-model notice: Apache-2.0 (`LICENSE-Qwen`). |
34
 
35
- ## Results
36
-
37
- ### JevBench: public-item development proxies
38
-
39
- | Model | Public-items proxy | Public hard (111) | Hard ECE | Probability TVD | Official JevBench v1.4 |
40
- |---|---:|---:|---:|---:|---:|
41
- | **blink-4b** | 76.5 | 80/111 | 0.067 | 0.226 | no official score published |
42
- | JevK5 v0.2.0 | 76.1 | 79/111 (own runtime: 82/111) | 0.068 | 0.220 | 62.0 |
43
- | Jev 1.13.0 | — | — | — | — | 63.3 |
44
-
45
- No official JevBench score for blink has been published. The JevBench numbers are public-item development proxies, not official scores, and claim no rank or parity. Official scoring requires held-out, judge, and sealed items.
46
-
47
- <details><summary>How to read the public-item numbers</summary>
48
-
49
- - On the public hard items, blink-4b scored 80/111 and JevK5 scored 79/111 in the same runtime; JevK5's own runtime reports 82/111. No hard-accuracy advantage is claimed. blink-4b's 95% Wilson interval is 0.631–0.796, before selection effects.
50
- - Speed comes from each row's serial run, applying JevBench's self-hosted adjustment.
51
- - Cost uses JevBench's 4B tariff ($0.03 per million input tokens) times measured tokens per decision; not a production bill.
52
-
53
- </details>
54
-
55
- ### Decision Index 0.2 (local run)
56
-
57
- Decision Index numbers are local runs of the official kit (commit 19ad28e on 2026-09-25), not leaderboard submissions. The 0.2 run is descriptive: known training exposure stays in the scores, with no leaderboard-style penalty, so it isn't ranked. Training included 281 MMLU-Pro test-partition questions, which contaminate the 0.2 MMLU-Pro score. "Without MMLU-Pro" is a sensitivity check, not a clean score.
58
-
59
- | Balanced skill | Balanced raw | Breadth skill | Without MMLU-Pro |
60
- |---:|---:|---:|---:|
61
- | 37.85 | 53.33 | 36.78 | 37.41 |
62
-
63
- <details><summary>Breakdown by area and added benchmarks</summary>
64
-
65
- | Area | Number of benchmarks | Skill | Raw |
66
- |---|---:|---:|---:|
67
- | Knowledge & Reasoning | 10 | 26.4 | 43.1 |
68
- | Language Understanding | 10 | 47.4 | 62.7 |
69
- | Retrieval & Classification | 7 | 36.8 | 54.7 |
70
- | Tools & Automation | 6 | 51.6 | 60.0 |
71
- | Arts & Human Taste | 7 | 27.2 | 46.1 |
72
-
73
- Benchmarks added in 0.2:
74
-
75
- | Benchmark | Metric | Requests | Answered | Raw | Skill |
76
- |---|---|---:|---:|---:|---:|
77
- | PhishNChips phishing decisions | accuracy | 2,000 | 2,000 | 63.6 | 27.3 |
78
- | MMLU-Pro | accuracy | 12,032 | 12,032 | 52.1 | 46.1 |
79
- | BBH fixed-option tasks | accuracy | 5,507 | 5,507 | 63.8 | 47.5 |
80
- | RAGTruth response-level hallucination | F1 on hallucinated class | 2,700 | 2,700 | 66.5 | 43.1 |
81
- | HoVer claim verification | accuracy | 4,000 | 4,000 | 63.1 | 26.2 |
82
- | When2Call MCQ | accuracy | 3,652 | 3,652 | 62.8 | 50.3 |
83
- | New Yorker caption matching | accuracy | 528 | 528 | 58.9 | 48.6 |
84
-
85
- - All 151,034 scoreable requests scored across 40 counted benchmarks in five equal areas.
86
- - Shared requests reuse 0.1 predictions; the 30,419 added requests ran with the frozen evaluated soup at temperature 1.0.
87
- - Point estimates only; no significance, calibration, or latency claims. Do not compare with 0.1 due to edition differences.
88
-
89
- Training-row text matches in added requests:
90
-
91
- | Training stage | Rows in the stage | Rows matching added-request text | From MMLU-Pro | From SuperGPQA | Other |
92
- |---|---:|---:|---:|---:|---:|
93
- | T3 | 23,156 | 138 | 131 | 6 | 1 |
94
- | T4 | 42,360 | 156 | 148 | 7 | 1 |
95
-
96
- Exact normalised strings >= 30 characters shared by training rows and added requests (excluding strings in >= 20 requests). Counts reflect training rows by stage and source, not unique test questions.
97
-
98
- </details>
99
-
100
- <details><summary>Decision Index 0.1 (archived edition)</summary>
101
-
102
- Local run of the archived 0.1 suite (132,422 requests across 37 benchmarks; 19 panel benchmarks averaged for headline index; comparison rows from 2026-09-22 leaderboard snapshot).
103
-
104
- | Model | Size class | Decision Index 0.1 | Skill | Breadth |
105
- |---|---|---:|---:|---:|
106
- | **blink-4b** | 4B | 52.12 | 36.04 | 34.18 |
107
- | Jev 1.13.0 | closed | 59.51 | 46.26 | 44.79 |
108
- | Jevfire | 27B | 55.74 | 40.86 | 39.45 |
109
- | JoshuaSP diffusiongemma (open-jev) | 26B-A4B | 55.56 | 40.84 | 39.19 |
110
- | Decider 35B-A3B | 35B-A3B | 54.34 | 39.37 | 37.99 |
111
- | Kev 9B | 9B | 50.48 | 32.96 | 30.54 |
112
- | Kev 4B | 4B | 47.43 | 28.86 | 25.67 |
113
-
114
- | Area | blink-4b | Jev 1.13.0 |
115
- |---|---:|---:|
116
- | Knowledge & Reasoning | 49.1 | 68.8 |
117
- | Language Understanding | 61.0 | 62.3 |
118
- | Retrieval & Classification | 30.1 | 37.0 |
119
- | Tools & Automation | 70.2 | 73.6 |
120
- | Arts & Human Judgment | 50.2 | 56.2 |
121
-
122
- Per benchmark scores (19 panel benchmarks, 0.1):
123
-
124
- | Area | Benchmark | blink-4b | Jev 1.13.0 |
125
- |---|---|---:|---:|
126
- | Knowledge | MMLU | 0.749 | 0.917 |
127
- | Knowledge | GPQA Diamond | 0.372 | 0.783 |
128
- | Knowledge | GSM8K | 0.579 | 0.799 |
129
- | Knowledge | CRUXEval | 0.472 | 0.730 |
130
- | Knowledge | CLadder | 0.637 | 0.726 |
131
- | Knowledge | ChessBench | 0.135 | 0.172 |
132
- | Language | ContractNLI | 0.761 | 0.717 |
133
- | Language | iSarcasmEval | 0.452 | 0.505 |
134
- | Language | VAST | 0.617 | 0.646 |
135
- | Retrieval | BRIGHT | 0.172 | 0.187 |
136
- | Retrieval | Amazon ESCI | 0.431 | 0.552 |
137
- | Tools | BFCL | 0.903 | 0.958 |
138
- | Tools | ToolRet | 0.412 | 0.450 |
139
- | Tools | RouterBench | 0.790 | 0.799 |
140
- | Arts | BPoMP | 0.847 | 0.906 |
141
- | Arts | Humicroedit | 0.605 | 0.619 |
142
- | Arts | POP909-CL | 0.034 | 0.181 |
143
- | Arts | cfcolor | 0.581 | 0.647 |
144
- | Arts | Habermas Machine | 0.443 | 0.459 |
145
-
146
- </details>
147
-
148
- ## Use
149
 
150
  ```python
151
  # pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
@@ -174,7 +63,7 @@ print(out["answers"]["intent"]["probabilities"])
174
 
175
  ## Run it as a server
176
 
177
- `serve.py` serves `POST /v1/systemone` and `GET /v1/models`, compatible with TypeSafe server-side Python and JavaScript SDKs (`TYPESAFE_BASE_URL`), JevBench's `typesafe` adapter, and Decision Index's `http` engine. `GET /healthz` reports startup checks. Optional authentication via `--api-key` or `BLINK_API_KEY` requires `Authorization: Bearer <key>` on API routes (returns 401 if missing or invalid; `/healthz` stays open). Error bodies provide details in `error` and `detail`; over-limit requests return 422 without truncation. Requests run one at a time by default, batching questions into single forward passes. Enable cross-request batching with `--batch-window-ms 5` (up to `--max-batch-requests 16`, `--max-queued-requests 64`). Excess queued requests return 529 with `Retry-After` (v1.1 returned 503). See the [API reference](https://thegovind.github.io/blink/api/) for details.
178
 
179
  ```sh
180
  pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
@@ -183,22 +72,13 @@ python blink-4b/serve.py --model ./blink-4b --port 8000
183
  # TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any
184
  ```
185
 
186
- Check health: `curl -s http://127.0.0.1:8000/healthz`.
187
-
188
- Run with Docker:
189
 
190
  ```sh
191
  cd blink-4b
192
  docker build -t blink-4b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-4b
193
  ```
194
 
195
- To enable cross-request batching:
196
-
197
- ```sh
198
- hf download thegovind/blink-4b serve.py blink.py --revision v1.4 --local-dir blink-4b
199
- python blink-4b/serve.py --model ./blink-4b --port 8000 --batch-window-ms 5
200
- ```
201
-
202
  ## Higher throughput (opt-in)
203
 
204
  `serve_vllm.py` (added in `v1.4`) is an opt-in, text-only server with higher throughput. `serve.py` stays the default.
@@ -226,79 +106,43 @@ python serve_vllm.py --model . --port 8000 --quantization auto --max-concurrency
226
 
227
  ## Screenshots (opt-in, self-hosted)
228
 
229
- Try [Screen click](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen), read by blink-mimo-9b in the Space. See [Computer use](https://thegovind.github.io/blink/computer-use/) for setup and links.
230
-
231
- Image input is off by default. Start `serve.py` with `--vision-tower Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, or set `BLINK_VISION_TOWER` to that value for direct `blink.py` calls. Cache the matching tower before serving a downloaded folder offline; image mode needs `torchvision==0.28.0`. The checkpoint's weights stay text-only.
232
-
233
- If loading only `blink.py` via `hf_hub_download`, also download `graft_keys.py` from the same repo and revision beside it.
234
 
235
- This is a self-hosted blink extension. TypeSafe's hosted Jev is text-only and its published API has no image field. Put a `data:image/png;base64,...` URI (JPEG or WebP also work) in a string in `state`, or send data URIs in a top-level `images` list. Image URLs are not fetched. `--image-layout first|inline` places image tokens; `first` is the default, not a switch that enables images. For renamed folders or Docker's `/blink`, use `--model-name blink-4b` or `BLINK_MODEL_NAME=blink-4b`; matching checks still apply.
236
 
237
- Defaults: 2 images per request, 8 MiB decoded per image, 20 million source pixels and 2,088,960 pixels after resizing. `--max-images`, `--max-image-bytes`, `--max-image-source-pixels` and `--max-image-pixels` can raise those limits. Invalid images return 422. Text-only requests keep the byte-identical pre-image prompt and numeric path; image requests alone add pixel and visual-token usage.
238
 
239
- These are **development readouts, not a benchmark result**. The `first` default came from an amended rule on previously seen development data. Earlier text-serving results on this card predate v1.3 image serving. Target accuracy with five offered choices:
240
 
241
- | Layout | ScreenSpot-v2 K5 target | GUIOdyssey K5 target | Mind2Web K5 given gold |
242
- |---|---:|---:|---:|
243
- | first (default) | 94.73% | 79.40% | 61.54% |
244
- | inline | 90.72% | 73.53% | 68.27% |
245
-
246
- Mind2Web shortlist recall is only 104/1,000: "given gold" excludes steps with no gold target offered. GUIOdyssey K5 target accuracy covers its labelled subset, not completed tasks.
247
 
248
- **Screenshot HTTP latency:** Measured on serving revision `f2414d5`, with default `first` layout and the reference PyTorch causal-conv1d fallback observed in retained server logs; 24 fresh requests per condition. These are not remeasurements on the released revision.
 
 
 
249
 
250
- | Image size | Questions | HTTP p50 | HTTP p95 |
251
- |---|---:|---:|---:|
252
- | 1280 x 720 | 1 | 158.23 ms | 169.27 ms |
253
- | 1920 x 1080 | 1 | 314.79 ms | 327.47 ms |
254
- | 1920 x 1080 | 4 | 1125.90 ms | 1170.53 ms |
255
 
256
- ## Details
257
 
258
- <details><summary>Architecture and training progression</summary>
259
 
260
  ![blink-4b network diagram: 32 decoder layers repeating 3 Gated DeltaNet layers then one full-attention layer (24 and 8 in total, hidden 2560), with full attention at 0-based layers 3, 7, … 31 as in the tensor names; LoRA rank 16 on every attention, Gated DeltaNet and MLP projection (32.5M parameters, merged after training); token embeddings, norms and lm_head frozen, with lm_head tied to the token embeddings, one matrix; the vision encoder and multi-token-prediction head removed; and the answer read from the offered option-letter rows of lm_head.](assets/blink-4b-network.png)
261
 
262
- | | What ships |
263
- |---|---|
264
- | Backbone | Qwen3.5-4B text model; 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 2560; 4,205,751,296 shipped text parameters |
265
- | Tuned | 32.5M LoRA parameters, merged before averaging |
266
- | Final weights | Uniform weight average ("soup") of T3, T4 step 300 and T4 final |
267
 
268
- **LoRA targets (rank 16, alpha 32, every language-model layer):** full-attention `q_proj`, `k_proj`, `v_proj`, `o_proj`; Gated DeltaNet `in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `out_proj`; and every MLP's `gate_proj`, `up_proj`, `down_proj`. Token embeddings, all norms and `lm_head` stayed frozen. The trained adapters were merged into the text weights.
269
 
270
- The vision encoder and multi-token-prediction (MTP) head were cut: 0 vision tensors, 0 MTP tensors.
271
-
272
- Readout takes next-token logits from offered option labels in `lm_head` (single tokens A–Z, then two-letter labels), computed in FP32 and softmaxed over offered letters. The rest of the vocabulary is ignored. These are option-conditional model probabilities, not certified chances of being right.
273
-
274
- Training used supervised fine-tuning with cross-entropy against target distributions: exact probabilities, teacher-verified probabilities, or one-hot labels, plus base-model KL anchor distributions on teacher-rejected prompts. Choice and yes/no options reshuffled each epoch; score levels maintained order. No RL or preference optimization.
275
 
276
  ![blink-4b post-training diagram: T3 (23,156 rows) and T4 (42,360 rows) broken down by data category feed supervised fine-tuning with cross-entropy over the offered option letters; both runs start from the base, and T3, T4 step 300 and T4 final are merged and averaged into release v1.0. Decision Index 0.1 full suite 52.12.](assets/blink-4b-post-training.png)
277
 
278
- Training progression across steps:
279
-
280
- | Step | DI 0.1 | Public hard | Hard ECE | Notes |
281
- |---|---:|---:|---:|---|
282
- | Qwen3.5-4B, zero-shot | 44.45 DI-S | 0.595 | 0.131 | Baseline before decision training. |
283
- | T1 (106.7k rows; lr 1e-4; 355 steps) | 53.15 DI-S | 0.559 | — | Public train splits and exact-probability items lifted DI-S but hurt hard items; NLI/classification did not transfer to long documents. |
284
- | T3 (23,156 question rows; lr 3e-5; 96 steps) | — | 0.649 | 0.089 | Restarted from base with worlds, exact probabilities, teacher-written docs, ~10% public replay, and 7.6% base anchors. |
285
- | T4 (42,360 question rows; lr 4e-5; 472 steps) | — | 0.712 (step 300); 0.676 (final) | — | Added judge-style items; earlier checkpoint performed better on hard cases. |
286
- | **Soup (T3 + T4 step 300 + T4 final)** | 52.12 full | 0.721 | 0.067 | Averaged three checkpoints for hard accuracy and calibration; shipped. |
287
-
288
- Negative results: distilling 27B answers into 4B did not help on hard items; a DI-focused 4B gained +0.4 on DI-S; mixing JevK5 weights into the soup did not help; Qwen3.5-9B with the T1 recipe scored 54.5 DI-S.
289
-
290
- </details>
291
-
292
- <details><summary>Training data and data licenses</summary>
293
-
294
- ### Training data mix
295
 
296
  | Stage | Question rows | Mix |
297
  |---|---:|---|
298
- | T3 | 23,156 | 11,352 decision worlds · 3,741 teacher-written question rows · 2,579 exact-probability worlds · 2,221 public-source (~10%) · 1,763 base-model anchors (7.6%) · 1,500 program-generated reasoning |
299
- | T4 | 42,360 | 12,000 decision worlds · 7,860 teacher-written question rows · 7,000 judge-style · 6,220 exact-probability worlds · 3,500 base-model anchors (8.3%) · 3,000 program-generated reasoning · 2,780 public-source |
300
-
301
- T4 judge-style includes 3,000 GSM8K-train solution checks, 2,500 Dolly-15k routing, and 1,500 program-answer checks. Teacher documents and typed questions were generated by Qwen3.8-27B and kept only when a blind re-solve by the same teacher agreed.
302
 
303
  ### Data sources and licences
304
 
@@ -317,18 +161,13 @@ T4 judge-style includes 3,000 GSM8K-train solution checks, 2,500 Dolly-15k routi
317
  | VAST, Humicroedit, OpenBookQA | None stated by source |
318
  | Code-generated worlds and teacher-written documents (Qwen3.8-27B) | See LICENSE.md |
319
 
320
- These are source-repository licences; they do not settle rights in every underlying text.
321
-
322
- </details>
323
 
324
- <details><summary>Evaluation notes and limits</summary>
325
 
326
- ### Evaluation notes
327
 
328
- - **Held-out and selection:** Candidate checkpoints and prompt format were selected on DI-S (3,000 requests). blink-4b was fixed using JevBench development proxies before its full-suite run. It scored 52.29 on the 129,422 requests outside DI-S. Public JevBench items (231 items) and JevK5's 65 hand-written hard items served as development selection sets; none were included in training. A one-time lockbox of 396 held-out authored items from domains unseen in any of the three checkpoints' training data (same generator families, not JevBench's sealed set) scored 0.861 accuracy and 0.026 ECE.
329
- - **Training overlap and audit:** Public train splits also used by the 0.1 index include ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, and GSM8K train split. ANLI and BANKING77 train splits were also used. Public sources included 281 MMLU-Pro test-partition questions and 2 GPQA extended-set questions. An audit of all question rows against the 0.1 suite and JevBench public items found no content matches (all 36 flags were the BANKING77 template); no public JevBench items or chess positions matched. The 13-word passage check didn't search long-document bodies or option text. Semantic or pretraining overlap can't be ruled out, and private JevBench items weren't available to check.
330
- - **Temperature:** A split held out from T4 suggested T = 0.82, with negligible gain. But 166 of its 401 items were in T3 training, so it isn't held out from the released average. Temperature 1.0 is retained without fitting.
331
- - **Limits:** English-centric (Arabic task A scored 0.313, task C pairs 0.645 on 0.1; broader multilingual ability is unestablished); does not chat or explain answers; text in state can sway answers; date arithmetic and long policies are the weakest cases. For `blink.py` and default `serve.py`: 255 options per choice, 2–10 score levels, 131,072 input tokens per question (longest evaluated prompt: 37,906 tokens) and 512 questions per request; over-limit requests get HTTP 422 with the reason, never truncated.
332
 
333
  </details>
334
 
@@ -336,4 +175,4 @@ These are source-repository licences; they do not settle rights in every underly
336
 
337
  Code: Apache-2.0. Weights: non-commercial research and evaluation only; see each model card's license.
338
 
339
- See [LICENSE.md](LICENSE.md) for the weight terms; the base model is Apache-2.0 (`LICENSE-Qwen`).
 
17
 
18
  # blink-4b
19
 
 
 
20
  ![blink-4b one-pass readout diagram: a state and typed choice, noul and score questions are rendered as evidence, criterion and lettered options; questions are batched and each batch is one forward pass, giving next-token logits at the answer position, and an FP32 softmax over the offered letters gives one probability per option. No text is generated.](assets/blink-4b-readout.png)
21
 
22
+ Small, fast decisions for routing and checks at volume. Send text or JSON `state` with `choice`, `noul` (yes/no), or `score` questions. Get a probability for every offered answer, not generated text. Each batch takes one forward pass; large requests can use several batches.
23
+
24
+ Use `choice` to route a request, `noul` for a yes/no check, or `score` for an ordered rating. The same call can ask several questions about a single state.
25
+
26
+ **Try it:** [Space](https://huggingface.co/spaces/thegovind/blink) · [Screen click](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen) · [Computer use](https://thegovind.github.io/blink/computer-use/) · [API](https://thegovind.github.io/blink/api/) · [Docs](https://thegovind.github.io/blink/) · [GitHub](https://github.com/thegovind/blink) · [blink-mimo-9b](https://huggingface.co/thegovind/blink-mimo-9b) · [blink-27b](https://huggingface.co/thegovind/blink-27b)
27
 
28
  ## At a glance
29
 
30
  | Attribute | Detail |
31
  |---|---|
32
+ | Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B), text weights only |
33
  | Weights size | 8.4 GB (bf16, 4,205,751,296 parameters) |
34
  | Revision | v1.4 (code revision; weights identical to v1.0) |
35
  | License | Weights: non-commercial research and evaluation only ([LICENSE.md](LICENSE.md)); code: Apache-2.0. Base-model notice: Apache-2.0 (`LICENSE-Qwen`). |
36
 
37
+ ## Quickstart
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38
 
39
  ```python
40
  # pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
 
63
 
64
  ## Run it as a server
65
 
66
+ `serve.py` serves `POST /v1/systemone`, `GET /v1/models`, and `GET /healthz`. Point TypeSafe's server-side Python or JavaScript SDKs at it with `TYPESAFE_BASE_URL`; text decisions use the same request and response fields as hosted Jev. Requests run one at a time by default; `--batch-window-ms 5` enables cross-request batching. The [API reference](https://thegovind.github.io/blink/api/) covers limits and errors.
67
 
68
  ```sh
69
  pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
 
72
  # TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any
73
  ```
74
 
75
+ Or use Docker from the downloaded folder:
 
 
76
 
77
  ```sh
78
  cd blink-4b
79
  docker build -t blink-4b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-4b
80
  ```
81
 
 
 
 
 
 
 
 
82
  ## Higher throughput (opt-in)
83
 
84
  `serve_vllm.py` (added in `v1.4`) is an opt-in, text-only server with higher throughput. `serve.py` stays the default.
 
106
 
107
  ## Screenshots (opt-in, self-hosted)
108
 
109
+ Image input is off by default. Start `serve.py` with `--vision-tower Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` to attach the matching tower. This is a self-hosted blink extension. TypeSafe's hosted Jev is text-only.
 
 
 
 
110
 
111
+ With `--vision-tower`, self-hosted blink borrows the pinned Qwen base model's vision encoder while its checkpoint stays text-only; [blink-mimo-9b](https://huggingface.co/thegovind/blink-mimo-9b) uses its own encoder and runs [Screen click](https://huggingface.co/spaces/thegovind/blink?tab=use-cases&case=screen).
112
 
113
+ Put a `data:image/png;base64,...` URI (JPEG and WebP data URIs work too) inside a `state` string, or pass data URIs in a top-level `images` list. Image URLs are never fetched. If loading only `blink.py` via `hf_hub_download`, also download `graft_keys.py` from the same repo and revision beside it.
114
 
115
+ See [Computer use](https://thegovind.github.io/blink/computer-use/) to self-host this model with images.
116
 
117
+ ## Results
 
 
 
 
 
118
 
119
+ | Local development readout | Result |
120
+ |---|---:|
121
+ | JevBench public-items proxy | 76.5; 80/111 hard; hard ECE 0.067 |
122
+ | Decision Index 0.2 balanced skill | 37.85 |
123
 
124
+ No official JevBench score for blink has been published. The JevBench numbers are public-item development proxies, not an official score, rank, or parity claim. Decision Index is a descriptive local run of the official kit, not a leaderboard submission; training overlap affects its scores.
 
 
 
 
125
 
126
+ <details><summary>Model details: architecture, training, data</summary>
127
 
128
+ ### Architecture and readout
129
 
130
  ![blink-4b network diagram: 32 decoder layers repeating 3 Gated DeltaNet layers then one full-attention layer (24 and 8 in total, hidden 2560), with full attention at 0-based layers 3, 7, … 31 as in the tensor names; LoRA rank 16 on every attention, Gated DeltaNet and MLP projection (32.5M parameters, merged after training); token embeddings, norms and lm_head frozen, with lm_head tied to the token embeddings, one matrix; the vision encoder and multi-token-prediction head removed; and the answer read from the offered option-letter rows of lm_head.](assets/blink-4b-network.png)
131
 
132
+ Qwen3.5-4B text backbone: 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 2560, and 4,205,751,296 shipped parameters. The training updated 32.5M LoRA parameters at rank 16, alpha 32: `q_proj`, `k_proj`, `v_proj`, `o_proj`; `in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `out_proj`; and `gate_proj`, `up_proj`, `down_proj`. The vision encoder and MTP head were removed: 0 vision tensors, 0 MTP tensors.
 
 
 
 
133
 
134
+ Readout softmaxes FP32 next-token logits over the offered option labels. These are option-conditional probabilities, not certified chances of being right.
135
 
136
+ ### Training
 
 
 
 
137
 
138
  ![blink-4b post-training diagram: T3 (23,156 rows) and T4 (42,360 rows) broken down by data category feed supervised fine-tuning with cross-entropy over the offered option letters; both runs start from the base, and T3, T4 step 300 and T4 final are merged and averaged into release v1.0. Decision Index 0.1 full suite 52.12.](assets/blink-4b-post-training.png)
139
 
140
+ Supervised fine-tuning on target distributions; no RL or preference optimization. T3 used lr 3e-5 for 96 steps; T4 used lr 4e-5 for 472 steps. The released weights average T3, T4 step 300, and T4 final.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
141
 
142
  | Stage | Question rows | Mix |
143
  |---|---:|---|
144
+ | T3 | 23,156 | 11,352 decision worlds · 3,741 teacher-written rows · 2,579 exact-probability worlds · 2,221 public-source · 1,763 base-model anchors · 1,500 program-generated reasoning |
145
+ | T4 | 42,360 | 12,000 decision worlds · 7,860 teacher-written rows · 7,000 judge-style · 6,220 exact-probability worlds · 3,500 base-model anchors · 3,000 program-generated reasoning · 2,780 public-source |
 
 
146
 
147
  ### Data sources and licences
148
 
 
161
  | VAST, Humicroedit, OpenBookQA | None stated by source |
162
  | Code-generated worlds and teacher-written documents (Qwen3.8-27B) | See LICENSE.md |
163
 
164
+ Source-repository licences do not settle rights in every underlying text.
 
 
165
 
166
+ ### Evaluation and limits
167
 
168
+ The archived Decision Index 0.1 local run scored 52.12; editions are not directly comparable. Public JevBench items were used for development selection, not training. Training included 281 MMLU-Pro test-partition questions, contaminating that Decision Index 0.2 component; semantic and pretraining overlap cannot be ruled out. English-centric, no chatting or explanations; text in `state` can sway an answer, especially on long policies and date arithmetic.
169
 
170
+ For `blink.py` and default `serve.py`: 255 options per choice, 2–10 score levels, 131,072 input tokens per question, and 512 questions per request; over-limit requests return 422 without truncation. For image placement use `--image-layout first|inline` (`first` is the default); for a renamed folder use `--model-name blink-4b`. Neither flag enables images on its own.
 
 
 
171
 
172
  </details>
173
 
 
175
 
176
  Code: Apache-2.0. Weights: non-commercial research and evaluation only; see each model card's license.
177
 
178
+ See [LICENSE.md](LICENSE.md) for weight terms; the Qwen base model is Apache-2.0 (`LICENSE-Qwen`).