Instructions to use thegovind/blink-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use thegovind/blink-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="thegovind/blink-4b")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("thegovind/blink-4b") model = AutoModelForCausalLM.from_pretrained("thegovind/blink-4b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download VLLM.md from thegovind/blink-4b: direct link, hf CLI and curl.
- Browser
- Download file 3.84 kB
-
https://huggingface.co/thegovind/blink-4b/resolve/main/VLLM.md
- Command line
-
hf download hf://thegovind/blink-4b/VLLM.md
-
curl -L -o VLLM.md https://huggingface.co/thegovind/blink-4b/resolve/main/VLLM.md
vLLM serving (opt-in)
serve_vllm.py serves blink-4b text-only through /v1/systemone, /v1/models and /healthz. serve.py stays the default; it can take screenshots when enabled.
From the model folder at revision v1.4:
Tested versions:
python -m pip install "torch==2.13.0" "transformers==5.17.0" \
"vllm==0.30.0" "compressed-tensors==0.17.0" \
"accelerate>=1.1.0" safetensors huggingface_hub
python -m pip install "flash-linear-attention==0.5.2"
python serve_vllm.py --model . --port 8000 \
--quantization auto --max-concurrency 32 --max-num-seqs 32 \
--max-num-batched-tokens 8192 --gpu-memory-utilization 0.85
Tested: vLLM 0.30.0; Transformers 5.17.0; Torch 2.13.0; compressed-tensors 0.17.0; flash-linear-attention 0.5.2. BF16 weights, FP32 offered-label head; chunked prefill on, prefix caching off.
Tested vLLM context: --max-model-len 32768; on this server, a question over 32,768 tokens returns 422. Default serve.py allows 131,072 tokens per question; another vLLM context size needs a new quality check.
Quality
The tables are from the earlier E3 study; capacity was not retimed and the other quality sets and c16 were not rerun. A release build re-passed the 231-item JevBench c1 quality check (80/111 hard, 199/231 total, hard ECE 0.069). Later changes touched only request checks, error handling and startup cleanup, not scoring.
| Task measure | vLLM result | Acceptance |
|---|---|---|
| JevBench public hard correct | 80/111 | at least 78/111 |
| JevBench public total correct | 199/231 | at least 196/231 |
| JevBench public hard top-label ECE (rounded) | c1 0.069; c16 first 0.069 / repeat 0.074 | at most ~0.077 on each |
| Web actions, 5 options (c1): agreement with the same model's FP32 answer | 499/500 | at least 489/500 |
| Web actions, 9 options (c1): agreement with the same model's FP32 answer | 494/500 | at least 487/500 |
| Decision Index 0.1 latency-tail sample (160 requests, 1,851 questions): agreement with the same model's FP32 answer | c1 1846/1851; c16 first 1847/1851 / repeat 1849/1851 | at least 1836/1851 on each |
Both web-action sets have 500 text-only questions each from public Multimodal-Mind2Web; ECE is rounded here, but checked unrounded. The Decision Index 0.1 latency-tail sample is 160 length-selected requests (1,851 questions) from reference-covered u1000, excluding source mismatches; it is not full u1000, DI-S or 0.2.
Capacity
| Measure | serve.py control |
serve_vllm.py |
|---|---|---|
| JevBench c1 p50 / p95 | 64 / 175 ms | 52 / 127 ms |
| Longest TypeSafe documents (16 requests), c1 p95 | 22.2 s | 16.4 s |
| JevBench c16 completed q/s | 12.2 | 34.7 |
| Decision Index 0.1 latency-tail sample (160 requests, 1,851 questions), c16 completed q/s | 21.0 | 46.5 |
| Peak device memory while scoring | 21.4 GiB | 68.0 GiB |
| Startup through health and warm score | first not captured; later control under 19 s | first cold start 180 s |
One self-hosted replica, loopback HTTP. c1 uses one request at a time; c16 uses 16 concurrent requests and reports completed questions/s. Peak memory includes reserved cache. Quality was checked on these samples, not on every Decision Index item. Agreement with the same model's FP32 answer is not gold accuracy or bit-for-bit parity. For more traffic, use separate warm replicas; recheck quality after changing flags or runtime. No official JevBench score for blink has been published.
Any request with a top-level images field (even []) or inline image data returns 422 ("this model reads text only"); the server reports accepts_images: false. --quantization auto keeps these BF16 weights. The INT8 builds that were tried did not pass the quality checks.
Code: Apache-2.0. Weights: non-commercial research and evaluation only; see LICENSE.md.