Text Classification
Transformers
Safetensors
English
qwen3_5_text
text-generation
decision-model
typed-decisions
one-pass
option-probabilities
Instructions to use thegovind/blink-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use thegovind/blink-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="thegovind/blink-4b")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("thegovind/blink-4b") model = AutoModelForCausalLM.from_pretrained("thegovind/blink-4b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download VLLM.md from thegovind/blink-4b: direct link, hf CLI and curl.
- Browser
- Download file 3.84 kB
-
https://huggingface.co/thegovind/blink-4b/resolve/f7f0e343e5f327d93a71e4cf2f0c56fd687c7665/VLLM.md
- Command line
-
hf download hf://thegovind/blink-4b@f7f0e343e5f327d93a71e4cf2f0c56fd687c7665/VLLM.md
-
curl -L -o VLLM.md https://huggingface.co/thegovind/blink-4b/resolve/f7f0e343e5f327d93a71e4cf2f0c56fd687c7665/VLLM.md
3.84 kB
| # vLLM serving (opt-in) | |
| `serve_vllm.py` serves blink-4b text-only through `/v1/systemone`, `/v1/models` and `/healthz`. `serve.py` stays the default; it can take screenshots when enabled. | |
| From the model folder at revision `v1.4`: | |
| Tested versions: | |
| ```sh | |
| python -m pip install "torch==2.13.0" "transformers==5.17.0" \ | |
| "vllm==0.30.0" "compressed-tensors==0.17.0" \ | |
| "accelerate>=1.1.0" safetensors huggingface_hub | |
| python -m pip install "flash-linear-attention==0.5.2" | |
| ``` | |
| ```sh | |
| python serve_vllm.py --model . --port 8000 \ | |
| --quantization auto --max-concurrency 32 --max-num-seqs 32 \ | |
| --max-num-batched-tokens 8192 --gpu-memory-utilization 0.85 | |
| ``` | |
| Tested: vLLM 0.30.0; Transformers 5.17.0; Torch 2.13.0; compressed-tensors 0.17.0; flash-linear-attention 0.5.2. BF16 weights, FP32 offered-label head; chunked prefill on, prefix caching off. | |
| Tested vLLM context: `--max-model-len 32768`; on this server, a question over 32,768 tokens returns `422`. Default `serve.py` allows 131,072 tokens per question; another vLLM context size needs a new quality check. | |
| ## Quality | |
| The tables are from the earlier E3 study; capacity was not retimed and the other quality sets and c16 were not rerun. A release build re-passed the 231-item JevBench c1 quality check (80/111 hard, 199/231 total, hard ECE 0.069). Later changes touched only request checks, error handling and startup cleanup, not scoring. | |
| | Task measure | vLLM result | Acceptance | | |
| | --- | ---: | ---: | | |
| | JevBench public hard correct | 80/111 | at least 78/111 | | |
| | JevBench public total correct | 199/231 | at least 196/231 | | |
| | JevBench public hard top-label ECE (rounded) | c1 0.069; c16 first 0.069 / repeat 0.074 | at most ~0.077 on each | | |
| | Web actions, 5 options (c1): agreement with the same model's FP32 answer | 499/500 | at least 489/500 | | |
| | Web actions, 9 options (c1): agreement with the same model's FP32 answer | 494/500 | at least 487/500 | | |
| | Decision Index 0.1 latency-tail sample (160 requests, 1,851 questions): agreement with the same model's FP32 answer | c1 1846/1851; c16 first 1847/1851 / repeat 1849/1851 | at least 1836/1851 on each | | |
| Both web-action sets have 500 text-only questions each from public Multimodal-Mind2Web; ECE is rounded here, but checked unrounded. The Decision Index 0.1 latency-tail sample is 160 length-selected requests (1,851 questions) from reference-covered u1000, excluding source mismatches; it is not full u1000, DI-S or 0.2. | |
| ## Capacity | |
| | Measure | `serve.py` control | `serve_vllm.py` | | |
| | --- | ---: | ---: | | |
| | JevBench c1 p50 / p95 | 64 / 175 ms | 52 / 127 ms | | |
| | Longest TypeSafe documents (16 requests), c1 p95 | 22.2 s | 16.4 s | | |
| | JevBench c16 completed q/s | 12.2 | 34.7 | | |
| | Decision Index 0.1 latency-tail sample (160 requests, 1,851 questions), c16 completed q/s | 21.0 | 46.5 | | |
| | Peak device memory while scoring | 21.4 GiB | 68.0 GiB | | |
| | Startup through health and warm score | first not captured; later control under 19 s | first cold start 180 s | | |
| One self-hosted replica, loopback HTTP. c1 uses one request at a time; c16 uses 16 concurrent requests and reports completed questions/s. Peak memory includes reserved cache. Quality was checked on these samples, not on every Decision Index item. Agreement with the same model's FP32 answer is not gold accuracy or bit-for-bit parity. For more traffic, use separate warm replicas; recheck quality after changing flags or runtime. No official JevBench score for blink has been published. | |
| Any request with a top-level `images` field (even `[]`) or inline image data returns `422` (`"this model reads text only"`); the server reports `accepts_images: false`. `--quantization auto` keeps these BF16 weights. The INT8 builds that were tried did not pass the quality checks. | |
| Code: Apache-2.0. Weights: non-commercial research and evaluation only; see `LICENSE.md`. | |