dewamade commited on
Commit
5a9322f
·
verified ·
1 Parent(s): abb4fe1

docs: serving profiles with KV pool / concurrency numbers

Browse files
Files changed (1) hide show
  1. README.md +11 -2
README.md CHANGED
@@ -35,8 +35,17 @@ To switch back to the INT8 embedding: restore the original
35
  `model-00002-of-00005.safetensors` + `config.json` from the main repo (keep a copy before
36
  overwriting, or re-download it).
37
 
38
- Serving is as in the main repo's README: the [HyperQwen](https://github.com/syv-ai/HyperQwen)
39
- vLLM fork single-user launcher (`CTX=long` profile, ~220K-token KV pool).
 
 
 
 
 
 
 
 
 
40
 
41
  ## License
42
 
 
35
  `model-00002-of-00005.safetensors` + `config.json` from the main repo (keep a copy before
36
  overwriting, or re-download it).
37
 
38
+ ## Serve
39
+
40
+ Serve the swapped model (e.g. a directory named `model-fast-emb4`) with the
41
+ [HyperQwen](https://github.com/syv-ai/HyperQwen) vLLM fork single-user launcher:
42
+
43
+ ```bash
44
+ MODEL=path/to/model-fast-emb4 PREFIX_CACHE=1 VISION=1 SPEC=mtp CTX=long REQ_METRICS=1 bash single-user/start_qwen.sh
45
+ ```
46
+
47
+ GPU KV cache size: **218,112 tokens** — maximum concurrency for 150,000 tokens per request:
48
+ **1.45x** (the INT4 embedding frees ~0.6 GB back to the KV pool vs the INT8 version).
49
 
50
  ## License
51