blink-4b / VLLM.md
thegovind's picture
Code revision v1.4: opt-in vLLM server (text only)
f7f0e34 verified
|
Raw History Blame Contribute Delete
3.84 kB

vLLM serving (opt-in)

serve_vllm.py serves blink-4b text-only through /v1/systemone, /v1/models and /healthz. serve.py stays the default; it can take screenshots when enabled.

From the model folder at revision v1.4:

Tested versions:

python -m pip install "torch==2.13.0" "transformers==5.17.0" \
  "vllm==0.30.0" "compressed-tensors==0.17.0" \
  "accelerate>=1.1.0" safetensors huggingface_hub
python -m pip install "flash-linear-attention==0.5.2"
python serve_vllm.py --model . --port 8000 \
  --quantization auto --max-concurrency 32 --max-num-seqs 32 \
  --max-num-batched-tokens 8192 --gpu-memory-utilization 0.85

Tested: vLLM 0.30.0; Transformers 5.17.0; Torch 2.13.0; compressed-tensors 0.17.0; flash-linear-attention 0.5.2. BF16 weights, FP32 offered-label head; chunked prefill on, prefix caching off.

Tested vLLM context: --max-model-len 32768; on this server, a question over 32,768 tokens returns 422. Default serve.py allows 131,072 tokens per question; another vLLM context size needs a new quality check.

Quality

The tables are from the earlier E3 study; capacity was not retimed and the other quality sets and c16 were not rerun. A release build re-passed the 231-item JevBench c1 quality check (80/111 hard, 199/231 total, hard ECE 0.069). Later changes touched only request checks, error handling and startup cleanup, not scoring.

Task measure vLLM result Acceptance
JevBench public hard correct 80/111 at least 78/111
JevBench public total correct 199/231 at least 196/231
JevBench public hard top-label ECE (rounded) c1 0.069; c16 first 0.069 / repeat 0.074 at most ~0.077 on each
Web actions, 5 options (c1): agreement with the same model's FP32 answer 499/500 at least 489/500
Web actions, 9 options (c1): agreement with the same model's FP32 answer 494/500 at least 487/500
Decision Index 0.1 latency-tail sample (160 requests, 1,851 questions): agreement with the same model's FP32 answer c1 1846/1851; c16 first 1847/1851 / repeat 1849/1851 at least 1836/1851 on each

Both web-action sets have 500 text-only questions each from public Multimodal-Mind2Web; ECE is rounded here, but checked unrounded. The Decision Index 0.1 latency-tail sample is 160 length-selected requests (1,851 questions) from reference-covered u1000, excluding source mismatches; it is not full u1000, DI-S or 0.2.

Capacity

Measure serve.py control serve_vllm.py
JevBench c1 p50 / p95 64 / 175 ms 52 / 127 ms
Longest TypeSafe documents (16 requests), c1 p95 22.2 s 16.4 s
JevBench c16 completed q/s 12.2 34.7
Decision Index 0.1 latency-tail sample (160 requests, 1,851 questions), c16 completed q/s 21.0 46.5
Peak device memory while scoring 21.4 GiB 68.0 GiB
Startup through health and warm score first not captured; later control under 19 s first cold start 180 s

One self-hosted replica, loopback HTTP. c1 uses one request at a time; c16 uses 16 concurrent requests and reports completed questions/s. Peak memory includes reserved cache. Quality was checked on these samples, not on every Decision Index item. Agreement with the same model's FP32 answer is not gold accuracy or bit-for-bit parity. For more traffic, use separate warm replicas; recheck quality after changing flags or runtime. No official JevBench score for blink has been published.

Any request with a top-level images field (even []) or inline image data returns 422 ("this model reads text only"); the server reports accepts_images: false. --quantization auto keeps these BF16 weights. The INT8 builds that were tried did not pass the quality checks.

Code: Apache-2.0. Weights: non-commercial research and evaluation only; see LICENSE.md.