thegovind commited on
Commit
a2e49b3
·
verified ·
1 Parent(s): 846ab75

Shorter card; link the API docs

Browse files
Files changed (1) hide show
  1. README.md +73 -175
README.md CHANGED
@@ -17,20 +17,20 @@ language:
17
 
18
  # blink-4b
19
 
20
- Send a text or JSON `state` and your questions: `choice` picks from up to 255 options, `noul` is yes/no,
21
- and `score` takes 2–10 ordered levels. Each question gets probabilities over its offered options from
22
- one forward pass, with no generated text. Long or large multi-question requests may use several batches.
23
 
24
- **Try it:** [Space demo](https://huggingface.co/spaces/thegovind/blink) — blink-4b and blink-mimo-9b run live ·
25
- [blink-4b](https://huggingface.co/thegovind/blink-4b) · [blink-27b](https://huggingface.co/thegovind/blink-27b) ·
26
- [blink-mimo-9b](https://huggingface.co/thegovind/blink-mimo-9b).
27
 
28
- [Source code](https://github.com/thegovind/blink) · [Docs](https://thegovind.github.io/blink/) · [blink-4b docs](https://thegovind.github.io/blink/models/blink-4b/).
29
 
30
- *Personal research release by thegovind, not an official product of any company. No affiliation with TypeSafe AI,
31
- Xiaomi, Alibaba Cloud or the Qwen team. Weights are for non-commercial research; see [Licence](#licence).*
32
 
33
- ![blink-4b one-pass readout diagram: a state and typed choice, noul and score questions are rendered as evidence, criterion and lettered options; questions are batched and each batch is one forward pass, giving next-token logits at the answer position, and an FP32 softmax over the offered letters gives one probability per option. No text is generated.](assets/blink-4b-readout.png)
 
 
 
 
 
34
 
35
  ## Results
36
 
@@ -42,27 +42,25 @@ Xiaomi, Alibaba Cloud or the Qwen team. Weights are for non-commercial research;
42
  | JevK5 v0.2.0 | 76.1 | 79/111 (own runtime: 82/111) | 0.068 | 0.220 | 62.0 |
43
  | Jev 1.13.0 | — | — | — | — | 63.3 |
44
 
45
- Both proxies use the same harness on public items. They aren't official JevBench scores or predictions: official scoring needs held-out, judge and sealed items. We claim no official score, rank or parity with Jev or JevK5.
46
 
47
  <details><summary>How to read the public-item numbers</summary>
48
 
49
- - On the public hard items, blink-4b got 80/111 and JevK5 got 79/111 in our runtime; JevK5's own runtime reports 82/111. We don't claim a hard-accuracy advantage. blink-4b's 95% Wilson interval is 0.631–0.796, before selection effects.
50
- - Speed comes from each row's own serial run. The proxy speed axis already applies JevBench's self-hosted adjustment.
51
  - Cost uses JevBench's 4B tariff ($0.03 per million input tokens) times measured tokens per decision; not a production bill.
52
 
53
  </details>
54
 
55
  ### Decision Index 0.2 (local run)
56
 
57
- We ran the full Decision Index 0.2 suite ourselves with the official scoring kit at commit 19ad28e on 2026-09-25. This is a descriptive run, not a leaderboard submission or accepted result. The kit's scorer does not apply the leaderboard's penalty for rows an entrant trained on, so known training exposure stays in these scores and they cannot be ranked against the leaderboard.
58
 
59
  | Balanced skill | Balanced raw | Breadth skill | Without MMLU-Pro |
60
  |---:|---:|---:|---:|
61
  | 37.85 | 53.33 | 36.78 | 37.41 |
62
 
63
- - Training included 281 MMLU-Pro test-partition questions, so the 0.2 MMLU-Pro score is contaminated. “Without MMLU-Pro” drops that benchmark but does not remove other training effects. Public train splits used in training are listed under “Training overlap” below.
64
-
65
- <details><summary>Extra tables and method</summary>
66
 
67
  | Area | Number of benchmarks | Skill | Raw |
68
  |---|---:|---:|---:|
@@ -72,7 +70,7 @@ We ran the full Decision Index 0.2 suite ourselves with the official scoring kit
72
  | Tools & Automation | 6 | 51.6 | 60.0 |
73
  | Arts & Human Taste | 7 | 27.2 | 46.1 |
74
 
75
- The seven benchmarks added in 0.2.
76
 
77
  | Benchmark | Metric | Requests | Answered | Raw | Skill |
78
  |---|---|---:|---:|---:|---:|
@@ -84,31 +82,28 @@ The seven benchmarks added in 0.2.
84
  | When2Call MCQ | accuracy | 3,652 | 3,652 | 62.8 | 50.3 |
85
  | New Yorker caption matching | accuracy | 528 | 528 | 58.9 | 48.6 |
86
 
87
- - All 151,034 of 151,034 scoreable requests scored. Of 44 scored benchmarks, 40 count toward the index across five equal areas.
88
- - Requests shared with 0.1 reuse the model's 0.1 predictions. We ran the 30,419 added requests with the same frozen evaluation setup as 0.1, the evaluated soup, which is the published weights, at temperature 1.0.
89
- - Balanced skill is the headline index. “Without MMLU-Pro” drops MMLU-Pro, averages the other nine Knowledge benchmarks, and keeps five equal areas. It is a sensitivity check, not a score free of training effects.
90
- - These are point estimates, with no significance, calibration, or latency claims. Do not compare them with 0.1 numbers because the editions differ.
91
 
92
- Training-row text matches in the added requests.
93
 
94
  | Training stage | Rows in the stage | Rows matching added-request text | From MMLU-Pro | From SuperGPQA | Other |
95
  |---|---:|---:|---:|---:|---:|
96
  | T3 | 23,156 | 138 | 131 | 6 | 1 |
97
  | T4 | 42,360 | 156 | 148 | 7 | 1 |
98
 
99
- We screened for exact normalised strings of at least 30 characters shared by training rows and added requests, ignoring strings found in 20 or more requests as templates. Counts are training rows by stage and source, not unique test questions. A matching option or passage need not be the same question, and a clean screen cannot rule out semantic or pretraining overlap.
100
-
101
- We did not produce the planned calibration read or a score without the DI-S selection sample.
102
 
103
  </details>
104
 
105
- ### Decision Index 0.1 (archived edition)
106
 
107
- **Full suite: blink-4b 52.12 vs Jev 1.13.0 59.51.**
108
 
109
  | Model | Size class | Decision Index 0.1 | Skill | Breadth |
110
  |---|---|---:|---:|---:|
111
- | **blink-4b** (this model) | 4B | 52.12 | 36.04 | 34.18 |
112
  | Jev 1.13.0 | closed | 59.51 | 46.26 | 44.79 |
113
  | Jevfire | 27B | 55.74 | 40.86 | 39.45 |
114
  | JoshuaSP diffusiongemma (open-jev) | 26B-A4B | 55.56 | 40.84 | 39.19 |
@@ -116,10 +111,6 @@ We did not produce the planned calibration read or a score without the DI-S sele
116
  | Kev 9B | 9B | 50.48 | 32.96 | 30.54 |
117
  | Kev 4B | 4B | 47.43 | 28.86 | 25.67 |
118
 
119
- We ran the complete archived 0.1 suite: 132,422 requests across 37 benchmarks. The headline index averages 19 panel benchmarks. Comparison rows use the 2026-09-22 leaderboard snapshot. We ran the official kit's scorer locally; these aren't leaderboard submissions. The live [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index) moved to 0.2 on 2026-09-24. Our local 0.2 run is in the section above.
120
-
121
- On the archived 0.1 board, the best open entry with 3.5–5B served parameters was Kev 4B at 47.43; this model scored 52.12.
122
-
123
  | Area | blink-4b | Jev 1.13.0 |
124
  |---|---:|---:|
125
  | Knowledge & Reasoning | 49.1 | 68.8 |
@@ -128,9 +119,9 @@ On the archived 0.1 board, the best open entry with 3.5–5B served parameters w
128
  | Tools & Automation | 70.2 | 73.6 |
129
  | Arts & Human Judgment | 50.2 | 56.2 |
130
 
131
- <details><summary>Per benchmark (19 panel benchmarks, 0.1)</summary>
132
 
133
- | Area | Benchmark | This model | Jev 1.13.0 |
134
  |---|---|---:|---:|
135
  | Knowledge | MMLU | 0.749 | 0.917 |
136
  | Knowledge | GPQA Diamond | 0.372 | 0.783 |
@@ -154,56 +145,6 @@ On the archived 0.1 board, the best open entry with 3.5–5B served parameters w
154
 
155
  </details>
156
 
157
- ## What we changed in the network
158
-
159
- ![blink-4b network diagram: 32 decoder layers repeating 3 Gated DeltaNet layers then one full-attention layer (24 and 8 in total, hidden 2560), with full attention at 0-based layers 3, 7, … 31 as in the tensor names; LoRA rank 16 on every attention, Gated DeltaNet and MLP projection (32.5M parameters, merged after training); token embeddings, norms and lm_head frozen, with lm_head tied to the token embeddings, one matrix; the vision encoder and multi-token-prediction head removed; and the answer read from the offered option-letter rows of lm_head.](assets/blink-4b-network.png)
160
-
161
- | | What ships |
162
- |---|---|
163
- | Backbone | Qwen3.5-4B text model; 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 2560; 4,205,751,296 shipped text parameters |
164
- | Tuned | 32.5M LoRA parameters, merged before averaging |
165
- | Final weights | Uniform weight average ("soup") of T3, T4 step 300 and T4 final |
166
-
167
- T3/T4 also used KL anchors to the base model: its distributions on prompts whose teacher answers failed verification.
168
-
169
- **LoRA targets (rank 16, alpha 32, every language-model layer):** full-attention `q_proj`, `k_proj`, `v_proj`, `o_proj`;
170
- Gated DeltaNet `in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `out_proj`; and every MLP's
171
- `gate_proj`, `up_proj`, `down_proj`. Token embeddings, all norms and `lm_head` stayed frozen.
172
- The trained adapters were merged into the text weights.
173
-
174
- Qwen3.5-4B is a vision-language model; this checkpoint ships only its text model. The vision encoder and multi-token-prediction (MTP) head were cut: 0 vision tensors, 0 MTP tensors.
175
-
176
- **The real cut is at readout:** no text generation. One prompt pass; next-token logits from only the
177
- offered option-label rows of `lm_head` (verified single tokens A–Z, then two-letter labels), computed in
178
- FP32 and softmaxed over those letters. The rest of the vocabulary is ignored.
179
-
180
- **Objective:** "calibration-oriented decision post-training" is plain supervised fine-tuning.
181
- Cross-entropy uses each row's target distribution: code-computed exact probabilities, probability
182
- targets in teacher-written questions kept after a blind re-solve by that same teacher agreed, and
183
- one-hot labels otherwise. For blink-4b, the base model's distributions on anchor rows are also targets. Choice and yes/no options and letter assignments are
184
- reshuffled each epoch; score levels keep their order. Jev's RLCD recipe isn't public; we didn't
185
- use or reproduce it. No RL or preference optimisation.
186
-
187
- ## The climb
188
-
189
- ![blink-4b post-training diagram: T3 (23,156 rows) and T4 (42,360 rows) broken down by data category feed supervised fine-tuning with cross-entropy over the offered option letters; both runs start from the base, and T3, T4 step 300 and T4 final are merged and averaged into release v1.0. Decision Index 0.1 full suite 52.12.](assets/blink-4b-post-training.png)
190
-
191
- DI-S is the 3,000-request sample. JevBench hard and ECE here are public-item numbers, not official scores.
192
-
193
- | Step | DI 0.1 | Public hard | Hard ECE | Why |
194
- |---|---:|---:|---:|---|
195
- | Qwen3.5-4B, zero-shot | 44.45 DI-S | 0.595 | 0.131 | Baseline before decision training. |
196
- | T1 (106.7k rows; lr 1e-4; 355 steps) | 53.15 DI-S | 0.559 | — | Public train splits plus exact-probability items lifted DI-S but hurt hard items; NLI/classification didn't transfer to long documents. |
197
- | T3 (23,156 question rows; lr 3e-5; 96 steps) | — | 0.649 | 0.089 | Restarted from base with worlds (including "can't tell"), exact probabilities, teacher-written docs, ~10% public replay and 7.6% base anchors. |
198
- | T4 (42,360 question rows; lr 4e-5; 472 steps) | — | 0.712 (step 300); 0.676 (final) | — | Added judge-style items; the earlier checkpoint did better on hard cases. |
199
- | **Soup (T3 + T4 step 300 + T4 final)** | 52.12 full | 0.721 | 0.067 | Averaged three checkpoints for hard accuracy and calibration; shipped. |
200
-
201
- **Tried, didn't keep:**
202
- - Distilling 27B answers into 4B didn't help on hard items.
203
- - A DI-focused 4B gained just +0.4 on DI-S.
204
- - Mixing JevK5 weights into the soup didn't help.
205
- - Qwen3.5-9B with the T1 recipe scored 54.5 DI-S, below our pre-set bar.
206
-
207
  ## Use
208
 
209
  ```python
@@ -233,13 +174,7 @@ print(out["answers"]["intent"]["probabilities"])
233
 
234
  ## Run it as a server
235
 
236
- `serve.py` handles TypeSafe's request and answer fields at `POST /v1/systemone` and lists its model at
237
- `GET /v1/models`. From server-side code, point TypeSafe's Python or JavaScript SDK at the server with
238
- `TYPESAFE_BASE_URL`;
239
- JevBench's stock `typesafe` adapter and the Decision Index kit's `http` engine still work unchanged.
240
- `GET /healthz` reports startup checks. `v1.2` changes code only; its weights are identical to `v1.0`.
241
-
242
- See the [wire-format reference](https://thegovind.github.io/blink/wire-format/) for full API details.
243
 
244
  ```sh
245
  pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
@@ -248,80 +183,68 @@ python blink-4b/serve.py --model ./blink-4b --port 8000
248
  # TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any
249
  ```
250
 
251
- In another terminal: `curl -s http://127.0.0.1:8000/healthz`.
252
 
253
- **Or use Docker** from the downloaded folder:
254
 
255
  ```sh
256
  cd blink-4b
257
  docker build -t blink-4b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-4b
258
  ```
259
 
260
- <details><summary>Health, limits and weights</summary>
261
-
262
- - `/healthz` reports `weights_verified` (weight, config and tokenizer files listed in `weights.sha256`
263
- are hashed before serving; a mismatch stops startup), `warmup.repeat_identical` (two matching warm-up
264
- answers), `kernels` (fast path or slower fallback without flash-linear-attention), `versions` and `hub_offline`.
265
- - Limits: 255 options per choice, 2–10 score levels, 131,072 input tokens per question and 512 questions
266
- per request. Over-limit requests get HTTP 422 with the reason; nothing is truncated.
267
- - `GET /v1/models` lists the one served model with a blank `release_date`. Every request uses that model
268
- regardless of its `model` field.
269
- - The server is open by default. Set `--api-key` or `BLINK_API_KEY` to require `Authorization: Bearer <key>` on both API
270
- routes. Missing or wrong keys get 401; `/healthz` stays open.
271
- - Error bodies put the reason in `error` and `detail`. Over-limit requests return 422, with nothing cut.
272
- - Requests run one at a time. Questions are batched; each batch takes one forward pass (large requests
273
- can take more than one). Serving the downloaded folder or Docker image enables Hugging Face offline
274
- mode before model loading (`hub_offline: true`). The server doesn't otherwise restrict network access.
275
- - Set `--batch-window-ms 5` to turn on cross-request batching with a 5 ms collection window, up to
276
- `--max-batch-requests` requests at a time, which defaults to 16. The window defaults to 0, so requests still
277
- run one at a time. `--max-queued-requests` lets up to 64 requests wait for a batch by default. Excess requests
278
- get HTTP 529 with `Retry-After`, so clients should retry. `v1.1` returned HTTP 503 for a full queue. On a
279
- 1,000-request Decision Index sample over HTTP, throughput rose about 20% with 4 concurrent clients and 24%
280
- with 16. One client saw no gain. Offline runs on long documents showed no meaningful gain. On the Decision
281
- Index sample, a set of long workflow documents, and the public TypeSafe cases, batched answers passed the same
282
- numerical-parity checks against an FP32 reference as one-at-a-time answers, covering argmax agreement and
283
- probability differences. TypeSafe documents were sent as JSON objects. A few near-tied answers can still flip.
284
- Batching arrived in `v1.1`. The current code revision is `v1.2`, with weights identical to `v1.0`. Update the
285
- two code files in an existing `v1.0` download, then restart with the flag:
286
-
287
- ```sh
288
  hf download thegovind/blink-4b serve.py blink.py --revision v1.2 --local-dir blink-4b
289
  python blink-4b/serve.py --model ./blink-4b --port 8000 --batch-window-ms 5
290
  ```
291
- - blink-4b weights are 8.4 GB in bf16. Long prompts need more memory.
292
 
293
- </details>
 
 
294
 
295
- <details><summary>Model and probability readout</summary>
296
 
297
- | | |
298
  |---|---|
299
- | Base | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (Apache-2.0), text weights only |
300
- | Adaptation | LoRA (r16, α32) on attention, Gated DeltaNet and MLP projections; uniform average of three merged checkpoints from two runs |
301
- | Readout | Verified single-token labels (A–Z, then two-letter labels); FP32 softmax of next-token logits / T over the offered labels |
302
- | Temperature | 1.0 (not fitted; see evaluation notes) |
303
- | Input limit | 131,072 tokens per question; longest evaluated prompt: 37,906 tokens; longer inputs are refused, never truncated |
304
- | Runtime | `blink.py` builds prompts and reads out probabilities for text decisions |
305
-
306
- These are option-conditional model probabilities, not certified chances of being right. Calibration can shift
307
- across tasks, domains and option sets. For `choice`, `confidence = (p_max − 1/K)/(1 − 1/K)` measures
308
- concentration, not correctness. For `score`, `score` is the expected 0-based level and `choice` is the
309
- most likely level.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
310
 
311
  </details>
312
 
313
- <details><summary>Training and data</summary>
314
 
315
- ### How it was trained
316
 
317
  | Stage | Question rows | Mix |
318
  |---|---:|---|
319
  | T3 | 23,156 | 11,352 decision worlds · 3,741 teacher-written question rows · 2,579 exact-probability worlds · 2,221 public-source (~10%) · 1,763 base-model anchors (7.6%) · 1,500 program-generated reasoning |
320
  | T4 | 42,360 | 12,000 decision worlds · 7,860 teacher-written question rows · 7,000 judge-style · 6,220 exact-probability worlds · 3,500 base-model anchors (8.3%) · 3,000 program-generated reasoning · 2,780 public-source |
321
 
322
- T4 judge-style: 3,000 GSM8K-train solution checks, 2,500 Dolly-15k routing, 1,500 program-answer checks.
323
-
324
- The released weights average three checkpoints fine-tuned from the base: T3 (lr 3e-5, 96 steps), T4 at step 300 and T4 at its final step 472 (lr 4e-5). The base-model anchor targets come from the base's own distributions on authored prompts whose teacher answers failed verification. Qwen3.8-27B wrote the teacher documents and their typed questions.
325
 
326
  ### Data sources and licences
327
 
@@ -338,9 +261,9 @@ The released weights average three checkpoints fine-tuned from the base: T3 (lr
338
  | SciQ | CC BY-NC 3.0 |
339
  | iSarcasmEval | MIT (upstream repository licence) |
340
  | VAST, Humicroedit, OpenBookQA | None stated by source |
341
- | Our code-generated worlds and teacher-written documents (Qwen3.8-27B) | See LICENSE.md |
342
 
343
- These are source-repository licences; they don't settle rights in every underlying text.
344
 
345
  </details>
346
 
@@ -348,38 +271,13 @@ These are source-repository licences; they don't settle rights in every underlyi
348
 
349
  ### Evaluation notes
350
 
351
- - **Scorer parity.** `blink.py` and the evaluation scorer matched prompts, labels and argmaxes on 296 requests (max |Δp| < 1e-7). From a fresh pinned install and a fresh Hub download, JevBench's stock `typesafe` adapter matched the development run on decision answers, probabilities and token usage across all 231 public items (easy 48/48, standard 71/72, hard 80/111); only the reported model identifier differed, so raw responses weren't byte-identical. The Decision Index kit's `http` engine scored DI-S 50.16 versus 50.08 in-process; its per-request timer on 1,000 random suite requests served serially measured a 66.4 ms median over HTTP versus 65.2 ms in-process.
352
- - **Selection.** We reused DI-S, the official kit's 3,000-request sample of the 0.1 suite, to pick the prompt format and candidate checkpoints. We fixed blink-4b using JevBench development proxies before its full-suite run. It scored 52.29 on the 129,422 requests outside DI-S.
353
- - **Training overlap.** Public train splits also used by the 0.1 index: ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, GSM8K (train split; solution-checking items). We also used ANLI and BANKING77 train splits; they're in the 0.1 suite but outside its index, and both are in the 0.2 panel. The audit below reports what was checked and any shared passages.
354
- - **Partitions.** Public-source data included training and development partitions, plus 281 MMLU-Pro test-partition questions and 2 GPQA extended-set questions outside GPQA Diamond; MMLU-Pro is not in the 0.1 suite but is in the 0.2 panel, so the 0.2 MMLU-Pro score is contaminated by those training questions and the 0.2 results are descriptive.
355
- - **Final-mixture audit.** Rechecked every question row (including teacher-written rows, plus base-model anchors) against the complete 0.1 suite (132,422 requests) and JevBench's 231 public items. The checks looked for exact matches of normalised strings of at least 30 characters in any field and shared 13-word passages in each row's question text (instructions, state.question, state.code). Strings or passages seen in 20 or more suite requests were treated as prompt templates and ignored. No public JevBench item matched under these checks; a separate position check found no shared chess positions.
356
- - **Suite overlap.** No content match with the 0.1 suite under these checks; all 36 flags were the fixed BANKING77 prompt template.
357
- - **Audit limits.** The 13-word passage check didn't search long-document bodies or option text. Semantic or pretraining overlap can't be ruled out, and private JevBench items weren't available to check.
358
- - **Generated reasoning.** Our programs computed the labels for CRUXEval-style code and CLadder-style causal questions; no items from those benchmarks were used. We didn't reuse the suite's GSM8K distractors.
359
- - **Teacher documents.** We kept Qwen3.8-27B's documents only if a fresh blind solve by that same teacher agreed with the answer. That's an agreement filter, not independent verification.
360
- - The repo ships no benchmark items, GPQA text, JevBench items or teacher traces.
361
- - **JevBench selection.** We scored all 231 public items on each candidate checkpoint and used JevK5's 65 hand-written hard items (Apache-2.0) as a second selection set. None went into training; these are development results.
362
- - **Temperature.** A split held out from T4 suggested T = 0.82, with negligible gain. But 166 of its 401 items were in T3 training, so it isn't held out from the released average. We kept T = 1.0 without fitting it.
363
- - **One-time lockbox.** We read held-out authored items from domains unseen in any of the three checkpoints' training data (same generator families, not JevBench's sealed set) once: accuracy 0.861 and ECE 0.026 over 396 items.
364
-
365
- ### Limits
366
-
367
- - English-centric. Training included Arabic iSarcasmEval rows; on the 0.1 suite, Arabic task A scored 0.313 and task C pairs 0.645. Broader multilingual performance hasn't been established.
368
- - Doesn't chat or explain answers.
369
- - Text in the state can sway the answer.
370
- - Date arithmetic and long policies are its weakest cases.
371
 
372
  </details>
373
 
374
- ## Licence
375
-
376
- Qwen/Qwen3.5-4B is Apache-2.0 (`LICENSE-Qwen`). The blink weights are for **non-commercial research and evaluation
377
- only** (`LICENSE.md`); commercial use isn't licensed. Training used non-commercial, share-alike and unlicensed
378
- sources (see the table above). It's unsettled whether their terms reach the weights, so check upstream terms too.
379
- `blink.py`, `serve.py` and the Dockerfile are Apache-2.0.
380
 
381
- <details><summary>Credits</summary>
382
-
383
- The Qwen team (base models). SemIf (MIT) for the evidence/criterion/options prompt layout. The Decision Index kit (MIT) and JevBench (MIT) for evaluation. JevK5 (Apache-2.0) for its hand-written hard items, used for evaluation only.
384
-
385
- </details>
 
17
 
18
  # blink-4b
19
 
20
+ Send a text or JSON `state` and typed questions: `choice` picks from up to 255 options, `noul` is yes/no, and `score` takes 2–10 ordered levels. Each question gets probabilities over its offered options from one forward pass, with no generated text. Long or large multi-question requests may use several batches.
 
 
21
 
22
+ ![blink-4b one-pass readout diagram: a state and typed choice, noul and score questions are rendered as evidence, criterion and lettered options; questions are batched and each batch is one forward pass, giving next-token logits at the answer position, and an FP32 softmax over the offered letters gives one probability per option. No text is generated.](assets/blink-4b-readout.png)
 
 
23
 
24
+ **Try it:** [Space demo](https://huggingface.co/spaces/thegovind/blink) · [blink-4b](https://huggingface.co/thegovind/blink-4b) · [blink-27b](https://huggingface.co/thegovind/blink-27b) · [blink-mimo-9b](https://huggingface.co/thegovind/blink-mimo-9b) · [Source code](https://github.com/thegovind/blink) · [Docs](https://thegovind.github.io/blink/) · [API](https://thegovind.github.io/blink/api/)
25
 
26
+ ## At a glance
 
27
 
28
+ | Attribute | Detail |
29
+ |---|---|
30
+ | Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (text model only; vision encoder and MTP head removed) |
31
+ | Weights size | 8.4 GB (bf16, 4,205,751,296 parameters) |
32
+ | Revision | v1.2 (code revision; weights identical to v1.0) |
33
+ | License | Non-commercial research only ([LICENSE.md](LICENSE.md)); base model Apache-2.0 (`LICENSE-Qwen`) |
34
 
35
  ## Results
36
 
 
42
  | JevK5 v0.2.0 | 76.1 | 79/111 (own runtime: 82/111) | 0.068 | 0.220 | 62.0 |
43
  | Jev 1.13.0 | — | — | — | — | 63.3 |
44
 
45
+ The JevBench numbers are public-item development proxies, not official scores, and claim no rank or parity. Official scoring requires held-out, judge, and sealed items.
46
 
47
  <details><summary>How to read the public-item numbers</summary>
48
 
49
+ - On the public hard items, blink-4b scored 80/111 and JevK5 scored 79/111 in the same runtime; JevK5's own runtime reports 82/111. No hard-accuracy advantage is claimed. blink-4b's 95% Wilson interval is 0.631–0.796, before selection effects.
50
+ - Speed comes from each row's serial run, applying JevBench's self-hosted adjustment.
51
  - Cost uses JevBench's 4B tariff ($0.03 per million input tokens) times measured tokens per decision; not a production bill.
52
 
53
  </details>
54
 
55
  ### Decision Index 0.2 (local run)
56
 
57
+ Decision Index numbers are local runs of the official kit (commit 19ad28e on 2026-09-25), not leaderboard submissions. The 0.2 run is descriptive: known training exposure stays in the scores, with no leaderboard-style penalty, so it isn't ranked. Training included 281 MMLU-Pro test-partition questions, which contaminate the 0.2 MMLU-Pro score. "Without MMLU-Pro" is a sensitivity check, not a clean score.
58
 
59
  | Balanced skill | Balanced raw | Breadth skill | Without MMLU-Pro |
60
  |---:|---:|---:|---:|
61
  | 37.85 | 53.33 | 36.78 | 37.41 |
62
 
63
+ <details><summary>Breakdown by area and added benchmarks</summary>
 
 
64
 
65
  | Area | Number of benchmarks | Skill | Raw |
66
  |---|---:|---:|---:|
 
70
  | Tools & Automation | 6 | 51.6 | 60.0 |
71
  | Arts & Human Taste | 7 | 27.2 | 46.1 |
72
 
73
+ Benchmarks added in 0.2:
74
 
75
  | Benchmark | Metric | Requests | Answered | Raw | Skill |
76
  |---|---|---:|---:|---:|---:|
 
82
  | When2Call MCQ | accuracy | 3,652 | 3,652 | 62.8 | 50.3 |
83
  | New Yorker caption matching | accuracy | 528 | 528 | 58.9 | 48.6 |
84
 
85
+ - All 151,034 scoreable requests scored across 40 counted benchmarks in five equal areas.
86
+ - Shared requests reuse 0.1 predictions; the 30,419 added requests ran with the frozen evaluated soup at temperature 1.0.
87
+ - Point estimates only; no significance, calibration, or latency claims. Do not compare with 0.1 due to edition differences.
 
88
 
89
+ Training-row text matches in added requests:
90
 
91
  | Training stage | Rows in the stage | Rows matching added-request text | From MMLU-Pro | From SuperGPQA | Other |
92
  |---|---:|---:|---:|---:|---:|
93
  | T3 | 23,156 | 138 | 131 | 6 | 1 |
94
  | T4 | 42,360 | 156 | 148 | 7 | 1 |
95
 
96
+ Exact normalised strings >= 30 characters shared by training rows and added requests (excluding strings in >= 20 requests). Counts reflect training rows by stage and source, not unique test questions.
 
 
97
 
98
  </details>
99
 
100
+ <details><summary>Decision Index 0.1 (archived edition)</summary>
101
 
102
+ Local run of the archived 0.1 suite (132,422 requests across 37 benchmarks; 19 panel benchmarks averaged for headline index; comparison rows from 2026-09-22 leaderboard snapshot).
103
 
104
  | Model | Size class | Decision Index 0.1 | Skill | Breadth |
105
  |---|---|---:|---:|---:|
106
+ | **blink-4b** | 4B | 52.12 | 36.04 | 34.18 |
107
  | Jev 1.13.0 | closed | 59.51 | 46.26 | 44.79 |
108
  | Jevfire | 27B | 55.74 | 40.86 | 39.45 |
109
  | JoshuaSP diffusiongemma (open-jev) | 26B-A4B | 55.56 | 40.84 | 39.19 |
 
111
  | Kev 9B | 9B | 50.48 | 32.96 | 30.54 |
112
  | Kev 4B | 4B | 47.43 | 28.86 | 25.67 |
113
 
 
 
 
 
114
  | Area | blink-4b | Jev 1.13.0 |
115
  |---|---:|---:|
116
  | Knowledge & Reasoning | 49.1 | 68.8 |
 
119
  | Tools & Automation | 70.2 | 73.6 |
120
  | Arts & Human Judgment | 50.2 | 56.2 |
121
 
122
+ Per benchmark scores (19 panel benchmarks, 0.1):
123
 
124
+ | Area | Benchmark | blink-4b | Jev 1.13.0 |
125
  |---|---|---:|---:|
126
  | Knowledge | MMLU | 0.749 | 0.917 |
127
  | Knowledge | GPQA Diamond | 0.372 | 0.783 |
 
145
 
146
  </details>
147
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
148
  ## Use
149
 
150
  ```python
 
174
 
175
  ## Run it as a server
176
 
177
+ `serve.py` serves `POST /v1/systemone` and `GET /v1/models`, compatible with TypeSafe server-side Python and JavaScript SDKs (`TYPESAFE_BASE_URL`), JevBench's `typesafe` adapter, and Decision Index's `http` engine. `GET /healthz` reports startup checks. Optional authentication via `--api-key` or `BLINK_API_KEY` requires `Authorization: Bearer <key>` on API routes (returns 401 if missing or invalid; `/healthz` stays open). Error bodies provide details in `error` and `detail`; over-limit requests return 422 without truncation. Requests run one at a time by default, batching questions into single forward passes. Enable cross-request batching with `--batch-window-ms 5` (up to `--max-batch-requests 16`, `--max-queued-requests 64`). Excess queued requests return 529 with `Retry-After` (v1.1 returned 503). See the [API reference](https://thegovind.github.io/blink/api/) for details.
 
 
 
 
 
 
178
 
179
  ```sh
180
  pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
 
183
  # TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any
184
  ```
185
 
186
+ Check health: `curl -s http://127.0.0.1:8000/healthz`.
187
 
188
+ Run with Docker:
189
 
190
  ```sh
191
  cd blink-4b
192
  docker build -t blink-4b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-4b
193
  ```
194
 
195
+ To enable cross-request batching:
196
+
197
+ ```sh
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
198
  hf download thegovind/blink-4b serve.py blink.py --revision v1.2 --local-dir blink-4b
199
  python blink-4b/serve.py --model ./blink-4b --port 8000 --batch-window-ms 5
200
  ```
 
201
 
202
+ ## Details
203
+
204
+ <details><summary>Architecture and training progression</summary>
205
 
206
+ ![blink-4b network diagram: 32 decoder layers repeating 3 Gated DeltaNet layers then one full-attention layer (24 and 8 in total, hidden 2560), with full attention at 0-based layers 3, 7, … 31 as in the tensor names; LoRA rank 16 on every attention, Gated DeltaNet and MLP projection (32.5M parameters, merged after training); token embeddings, norms and lm_head frozen, with lm_head tied to the token embeddings, one matrix; the vision encoder and multi-token-prediction head removed; and the answer read from the offered option-letter rows of lm_head.](assets/blink-4b-network.png)
207
 
208
+ | | What ships |
209
  |---|---|
210
+ | Backbone | Qwen3.5-4B text model; 32 decoder layers (24 Gated DeltaNet, 8 full-attention), hidden 2560; 4,205,751,296 shipped text parameters |
211
+ | Tuned | 32.5M LoRA parameters, merged before averaging |
212
+ | Final weights | Uniform weight average ("soup") of T3, T4 step 300 and T4 final |
213
+
214
+ **LoRA targets (rank 16, alpha 32, every language-model layer):** full-attention `q_proj`, `k_proj`, `v_proj`, `o_proj`; Gated DeltaNet `in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `out_proj`; and every MLP's `gate_proj`, `up_proj`, `down_proj`. Token embeddings, all norms and `lm_head` stayed frozen. The trained adapters were merged into the text weights.
215
+
216
+ The vision encoder and multi-token-prediction (MTP) head were cut: 0 vision tensors, 0 MTP tensors.
217
+
218
+ Readout takes next-token logits from offered option labels in `lm_head` (single tokens A–Z, then two-letter labels), computed in FP32 and softmaxed over offered letters. The rest of the vocabulary is ignored. These are option-conditional model probabilities, not certified chances of being right.
219
+
220
+ Training used supervised fine-tuning with cross-entropy against target distributions: exact probabilities, teacher-verified probabilities, or one-hot labels, plus base-model KL anchor distributions on teacher-rejected prompts. Choice and yes/no options reshuffled each epoch; score levels maintained order. No RL or preference optimization.
221
+
222
+ ![blink-4b post-training diagram: T3 (23,156 rows) and T4 (42,360 rows) broken down by data category feed supervised fine-tuning with cross-entropy over the offered option letters; both runs start from the base, and T3, T4 step 300 and T4 final are merged and averaged into release v1.0. Decision Index 0.1 full suite 52.12.](assets/blink-4b-post-training.png)
223
+
224
+ Training progression across steps:
225
+
226
+ | Step | DI 0.1 | Public hard | Hard ECE | Notes |
227
+ |---|---:|---:|---:|---|
228
+ | Qwen3.5-4B, zero-shot | 44.45 DI-S | 0.595 | 0.131 | Baseline before decision training. |
229
+ | T1 (106.7k rows; lr 1e-4; 355 steps) | 53.15 DI-S | 0.559 | — | Public train splits and exact-probability items lifted DI-S but hurt hard items; NLI/classification did not transfer to long documents. |
230
+ | T3 (23,156 question rows; lr 3e-5; 96 steps) | — | 0.649 | 0.089 | Restarted from base with worlds, exact probabilities, teacher-written docs, ~10% public replay, and 7.6% base anchors. |
231
+ | T4 (42,360 question rows; lr 4e-5; 472 steps) | — | 0.712 (step 300); 0.676 (final) | — | Added judge-style items; earlier checkpoint performed better on hard cases. |
232
+ | **Soup (T3 + T4 step 300 + T4 final)** | 52.12 full | 0.721 | 0.067 | Averaged three checkpoints for hard accuracy and calibration; shipped. |
233
+
234
+ Negative results: distilling 27B answers into 4B did not help on hard items; a DI-focused 4B gained +0.4 on DI-S; mixing JevK5 weights into the soup did not help; Qwen3.5-9B with the T1 recipe scored 54.5 DI-S.
235
 
236
  </details>
237
 
238
+ <details><summary>Training data and data licenses</summary>
239
 
240
+ ### Training data mix
241
 
242
  | Stage | Question rows | Mix |
243
  |---|---:|---|
244
  | T3 | 23,156 | 11,352 decision worlds · 3,741 teacher-written question rows · 2,579 exact-probability worlds · 2,221 public-source (~10%) · 1,763 base-model anchors (7.6%) · 1,500 program-generated reasoning |
245
  | T4 | 42,360 | 12,000 decision worlds · 7,860 teacher-written question rows · 7,000 judge-style · 6,220 exact-probability worlds · 3,500 base-model anchors (8.3%) · 3,000 program-generated reasoning · 2,780 public-source |
246
 
247
+ T4 judge-style includes 3,000 GSM8K-train solution checks, 2,500 Dolly-15k routing, and 1,500 program-answer checks. Teacher documents and typed questions were generated by Qwen3.8-27B and kept only when a blind re-solve by the same teacher agreed.
 
 
248
 
249
  ### Data sources and licences
250
 
 
261
  | SciQ | CC BY-NC 3.0 |
262
  | iSarcasmEval | MIT (upstream repository licence) |
263
  | VAST, Humicroedit, OpenBookQA | None stated by source |
264
+ | Code-generated worlds and teacher-written documents (Qwen3.8-27B) | See LICENSE.md |
265
 
266
+ These are source-repository licences; they do not settle rights in every underlying text.
267
 
268
  </details>
269
 
 
271
 
272
  ### Evaluation notes
273
 
274
+ - **Held-out and selection:** Candidate checkpoints and prompt format were selected on DI-S (3,000 requests). blink-4b was fixed using JevBench development proxies before its full-suite run. It scored 52.29 on the 129,422 requests outside DI-S. Public JevBench items (231 items) and JevK5's 65 hand-written hard items served as development selection sets; none were included in training. A one-time lockbox of 396 held-out authored items from domains unseen in any of the three checkpoints' training data (same generator families, not JevBench's sealed set) scored 0.861 accuracy and 0.026 ECE.
275
+ - **Training overlap and audit:** Public train splits also used by the 0.1 index include ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, and GSM8K train split. ANLI and BANKING77 train splits were also used. Public sources included 281 MMLU-Pro test-partition questions and 2 GPQA extended-set questions. An audit of all question rows against the 0.1 suite and JevBench public items found no content matches (all 36 flags were the BANKING77 template); no public JevBench items or chess positions matched. The 13-word passage check didn't search long-document bodies or option text. Semantic or pretraining overlap can't be ruled out, and private JevBench items weren't available to check.
276
+ - **Temperature:** A split held out from T4 suggested T = 0.82, with negligible gain. But 166 of its 401 items were in T3 training, so it isn't held out from the released average. Temperature 1.0 is retained without fitting.
277
+ - **Limits:** English-centric (Arabic task A scored 0.313, task C pairs 0.645 on 0.1; broader multilingual ability is unestablished); does not chat or explain answers; text in state can sway answers; date arithmetic and long policies are the weakest cases. Limits: 255 options per choice, 2–10 score levels, 131,072 input tokens per question (longest evaluated prompt: 37,906 tokens) and 512 questions per request; over-limit requests get HTTP 422 with the reason, never truncated.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
278
 
279
  </details>
280
 
281
+ ## License
 
 
 
 
 
282
 
283
+ Code (`blink.py`, `serve.py`, Dockerfile): Apache-2.0. Weights: non-commercial research only; see [LICENSE.md](LICENSE.md). Base model: Apache-2.0 (`LICENSE-Qwen`).