# vLLM serving (opt-in) `serve_vllm.py` serves blink-4b text-only through `/v1/systemone`, `/v1/models` and `/healthz`. `serve.py` stays the default; it can take screenshots when enabled. From the model folder at revision `v1.4`: Tested versions: ```sh python -m pip install "torch==2.13.0" "transformers==5.17.0" \ "vllm==0.30.0" "compressed-tensors==0.17.0" \ "accelerate>=1.1.0" safetensors huggingface_hub python -m pip install "flash-linear-attention==0.5.2" ``` ```sh python serve_vllm.py --model . --port 8000 \ --quantization auto --max-concurrency 32 --max-num-seqs 32 \ --max-num-batched-tokens 8192 --gpu-memory-utilization 0.85 ``` Tested: vLLM 0.30.0; Transformers 5.17.0; Torch 2.13.0; compressed-tensors 0.17.0; flash-linear-attention 0.5.2. BF16 weights, FP32 offered-label head; chunked prefill on, prefix caching off. Tested vLLM context: `--max-model-len 32768`; on this server, a question over 32,768 tokens returns `422`. Default `serve.py` allows 131,072 tokens per question; another vLLM context size needs a new quality check. ## Quality The tables are from the earlier E3 study; capacity was not retimed and the other quality sets and c16 were not rerun. A release build re-passed the 231-item JevBench c1 quality check (80/111 hard, 199/231 total, hard ECE 0.069). Later changes touched only request checks, error handling and startup cleanup, not scoring. | Task measure | vLLM result | Acceptance | | --- | ---: | ---: | | JevBench public hard correct | 80/111 | at least 78/111 | | JevBench public total correct | 199/231 | at least 196/231 | | JevBench public hard top-label ECE (rounded) | c1 0.069; c16 first 0.069 / repeat 0.074 | at most ~0.077 on each | | Web actions, 5 options (c1): agreement with the same model's FP32 answer | 499/500 | at least 489/500 | | Web actions, 9 options (c1): agreement with the same model's FP32 answer | 494/500 | at least 487/500 | | Decision Index 0.1 latency-tail sample (160 requests, 1,851 questions): agreement with the same model's FP32 answer | c1 1846/1851; c16 first 1847/1851 / repeat 1849/1851 | at least 1836/1851 on each | Both web-action sets have 500 text-only questions each from public Multimodal-Mind2Web; ECE is rounded here, but checked unrounded. The Decision Index 0.1 latency-tail sample is 160 length-selected requests (1,851 questions) from reference-covered u1000, excluding source mismatches; it is not full u1000, DI-S or 0.2. ## Capacity | Measure | `serve.py` control | `serve_vllm.py` | | --- | ---: | ---: | | JevBench c1 p50 / p95 | 64 / 175 ms | 52 / 127 ms | | Longest TypeSafe documents (16 requests), c1 p95 | 22.2 s | 16.4 s | | JevBench c16 completed q/s | 12.2 | 34.7 | | Decision Index 0.1 latency-tail sample (160 requests, 1,851 questions), c16 completed q/s | 21.0 | 46.5 | | Peak device memory while scoring | 21.4 GiB | 68.0 GiB | | Startup through health and warm score | first not captured; later control under 19 s | first cold start 180 s | One self-hosted replica, loopback HTTP. c1 uses one request at a time; c16 uses 16 concurrent requests and reports completed questions/s. Peak memory includes reserved cache. Quality was checked on these samples, not on every Decision Index item. Agreement with the same model's FP32 answer is not gold accuracy or bit-for-bit parity. For more traffic, use separate warm replicas; recheck quality after changing flags or runtime. No official JevBench score for blink has been published. Any request with a top-level `images` field (even `[]`) or inline image data returns `422` (`"this model reads text only"`); the server reports `accepts_images: false`. `--quantization auto` keeps these BF16 weights. The INT8 builds that were tried did not pass the quality checks. Code: Apache-2.0. Weights: non-commercial research and evaluation only; see `LICENSE.md`.