docs: serving profiles with KV pool / concurrency numbers
Browse files
README.md
CHANGED
|
@@ -35,8 +35,17 @@ To switch back to the INT8 embedding: restore the original
|
|
| 35 |
`model-00002-of-00005.safetensors` + `config.json` from the main repo (keep a copy before
|
| 36 |
overwriting, or re-download it).
|
| 37 |
|
| 38 |
-
|
| 39 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
|
| 41 |
## License
|
| 42 |
|
|
|
|
| 35 |
`model-00002-of-00005.safetensors` + `config.json` from the main repo (keep a copy before
|
| 36 |
overwriting, or re-download it).
|
| 37 |
|
| 38 |
+
## Serve
|
| 39 |
+
|
| 40 |
+
Serve the swapped model (e.g. a directory named `model-fast-emb4`) with the
|
| 41 |
+
[HyperQwen](https://github.com/syv-ai/HyperQwen) vLLM fork single-user launcher:
|
| 42 |
+
|
| 43 |
+
```bash
|
| 44 |
+
MODEL=path/to/model-fast-emb4 PREFIX_CACHE=1 VISION=1 SPEC=mtp CTX=long REQ_METRICS=1 bash single-user/start_qwen.sh
|
| 45 |
+
```
|
| 46 |
+
|
| 47 |
+
GPU KV cache size: **218,112 tokens** — maximum concurrency for 150,000 tokens per request:
|
| 48 |
+
**1.45x** (the INT4 embedding frees ~0.6 GB back to the KV pool vs the INT8 version).
|
| 49 |
|
| 50 |
## License
|
| 51 |
|