Default max-num-seqs=16, fp8-KV 128K serve example, measured concurrency (~10 full-128K sessions)
Browse files- README.md +8 -2
- USAGE.md +4 -3
- compose.yaml +1 -1
- entrypoint.sh +1 -1
README.md
CHANGED
|
@@ -86,13 +86,19 @@ answer in `content`. Omit it if you want the raw text (think tags included).
|
|
| 86 |
CUDA_VISIBLE_DEVICES=0 vllm serve sakamakismile/LFM2.5-8B-A1B-NVFP4 \
|
| 87 |
--served-model-name lfm25-8b-a1b \
|
| 88 |
--quantization modelopt \
|
| 89 |
-
--
|
| 90 |
-
--max-
|
|
|
|
| 91 |
--gpu-memory-utilization 0.90 \
|
| 92 |
--reasoning-parser deepseek_r1 \
|
| 93 |
--port 8000
|
| 94 |
```
|
| 95 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 96 |
| flag | what it does for *this* model |
|
| 97 |
|---|---|
|
| 98 |
| `--quantization modelopt` | **required** β reads `hf_quant_config.json` (NVFP4). Omit it and weights load as raw uint8 β garbage. |
|
|
|
|
| 86 |
CUDA_VISIBLE_DEVICES=0 vllm serve sakamakismile/LFM2.5-8B-A1B-NVFP4 \
|
| 87 |
--served-model-name lfm25-8b-a1b \
|
| 88 |
--quantization modelopt \
|
| 89 |
+
--kv-cache-dtype fp8 \
|
| 90 |
+
--max-model-len 128000 \
|
| 91 |
+
--max-num-seqs 16 \
|
| 92 |
--gpu-memory-utilization 0.90 \
|
| 93 |
--reasoning-parser deepseek_r1 \
|
| 94 |
--port 8000
|
| 95 |
```
|
| 96 |
|
| 97 |
+
**Concurrency on one 16 GB card (measured):** with `--kv-cache-dtype fp8` the KV pool is
|
| 98 |
+
**1.38 M tokens** β **~10 sessions each at a full 128 K context** (bf16 KV β ~5). With
|
| 99 |
+
paged KV and shorter prompts you can serve far more β `--max-num-seqs 16` is a good
|
| 100 |
+
default; pin to `10` to guarantee every slot at full 128 K.
|
| 101 |
+
|
| 102 |
| flag | what it does for *this* model |
|
| 103 |
|---|---|
|
| 104 |
| `--quantization modelopt` | **required** β reads `hf_quant_config.json` (NVFP4). Omit it and weights load as raw uint8 β garbage. |
|
USAGE.md
CHANGED
|
@@ -11,8 +11,9 @@ Tested: vLLM 0.21, CUDA 12.8/13.x, SM120. ~117 tok/s single, ~326 tok/s @4 concu
|
|
| 11 |
CUDA_VISIBLE_DEVICES=0 vllm serve /path/to/LFM2.5-8B-A1B-NVFP4 \
|
| 12 |
--served-model-name lfm25-8b-a1b \
|
| 13 |
--quantization modelopt \
|
| 14 |
-
--
|
| 15 |
-
--max-
|
|
|
|
| 16 |
--gpu-memory-utilization 0.90 \
|
| 17 |
--reasoning-parser deepseek_r1 \
|
| 18 |
--port 8000
|
|
@@ -34,7 +35,7 @@ curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -
|
|
| 34 |
|---|---|---|
|
| 35 |
| `--quantization modelopt` | **required** | tells vLLM the checkpoint is NVFP4 (`hf_quant_config.json`). Without it vLLM treats weights as raw uint8 β garbage. |
|
| 36 |
| `--max-model-len` | `32768` (up to `128000`) | context window. KV is cheap here (only 6 of 24 layers carry KV), so you can go long or run many sessions. |
|
| 37 |
-
| `--max-num-seqs` | `
|
| 38 |
| `--gpu-memory-utilization` | `0.90` | fraction of VRAM vLLM may use. 0.90 is safe on a clean 16 GB card; lower if the card is shared. |
|
| 39 |
| `--reasoning-parser deepseek_r1` | recommended | LFM2.5 is a **reasoning model** emitting `<think>β¦</think>` then the answer. This parser puts the CoT in `reasoning_content` and the answer in `content`. Omit it to get the raw text (think tags included). |
|
| 40 |
| `--kv-cache-dtype fp8` | optional | ~doubles KV capacity (β ~1.3 M tokens) for even more concurrency; tiny quality cost. Default (omit) keeps bf16 KV. |
|
|
|
|
| 11 |
CUDA_VISIBLE_DEVICES=0 vllm serve /path/to/LFM2.5-8B-A1B-NVFP4 \
|
| 12 |
--served-model-name lfm25-8b-a1b \
|
| 13 |
--quantization modelopt \
|
| 14 |
+
--kv-cache-dtype fp8 \
|
| 15 |
+
--max-model-len 128000 \
|
| 16 |
+
--max-num-seqs 16 \
|
| 17 |
--gpu-memory-utilization 0.90 \
|
| 18 |
--reasoning-parser deepseek_r1 \
|
| 19 |
--port 8000
|
|
|
|
| 35 |
|---|---|---|
|
| 36 |
| `--quantization modelopt` | **required** | tells vLLM the checkpoint is NVFP4 (`hf_quant_config.json`). Without it vLLM treats weights as raw uint8 β garbage. |
|
| 37 |
| `--max-model-len` | `32768` (up to `128000`) | context window. KV is cheap here (only 6 of 24 layers carry KV), so you can go long or run many sessions. |
|
| 38 |
+
| `--max-num-seqs` | `16` | concurrent sequences. Measured on a 16 GB card: with `--kv-cache-dtype fp8` the KV pool is **1.38 M tokens** β **~10 sessions each at a FULL 128 K context** (bf16 KV β ~5). 16 leaves headroom for shorter sessions via paged KV; pin to `10` if you want every slot guaranteed at full 128 K. |
|
| 39 |
| `--gpu-memory-utilization` | `0.90` | fraction of VRAM vLLM may use. 0.90 is safe on a clean 16 GB card; lower if the card is shared. |
|
| 40 |
| `--reasoning-parser deepseek_r1` | recommended | LFM2.5 is a **reasoning model** emitting `<think>β¦</think>` then the answer. This parser puts the CoT in `reasoning_content` and the answer in `content`. Omit it to get the raw text (think tags included). |
|
| 41 |
| `--kv-cache-dtype fp8` | optional | ~doubles KV capacity (β ~1.3 M tokens) for even more concurrency; tiny quality cost. Default (omit) keeps bf16 KV. |
|
compose.yaml
CHANGED
|
@@ -13,7 +13,7 @@ services:
|
|
| 13 |
- MODEL_DIR=/model
|
| 14 |
- PORT=8000
|
| 15 |
- MAX_MODEL_LEN=${MAX_MODEL_LEN:-32768}
|
| 16 |
-
- MAX_NUM_SEQS=${MAX_NUM_SEQS:-
|
| 17 |
- GPU_MEM_UTIL=${GPU_MEM_UTIL:-0.90}
|
| 18 |
- KV_CACHE_DTYPE=${KV_CACHE_DTYPE:-auto}
|
| 19 |
- TP_SIZE=${TP_SIZE:-1}
|
|
|
|
| 13 |
- MODEL_DIR=/model
|
| 14 |
- PORT=8000
|
| 15 |
- MAX_MODEL_LEN=${MAX_MODEL_LEN:-32768}
|
| 16 |
+
- MAX_NUM_SEQS=${MAX_NUM_SEQS:-16}
|
| 17 |
- GPU_MEM_UTIL=${GPU_MEM_UTIL:-0.90}
|
| 18 |
- KV_CACHE_DTYPE=${KV_CACHE_DTYPE:-auto}
|
| 19 |
- TP_SIZE=${TP_SIZE:-1}
|
entrypoint.sh
CHANGED
|
@@ -8,7 +8,7 @@ HOST="${HOST:-0.0.0.0}"
|
|
| 8 |
|
| 9 |
# --- tunables (all env-overridable) ---
|
| 10 |
MAX_MODEL_LEN="${MAX_MODEL_LEN:-32768}" # up to 128000
|
| 11 |
-
MAX_NUM_SEQS="${MAX_NUM_SEQS:-
|
| 12 |
GPU_MEM_UTIL="${GPU_MEM_UTIL:-0.90}"
|
| 13 |
KV_CACHE_DTYPE="${KV_CACHE_DTYPE:-auto}" # set "fp8" to ~double KV capacity
|
| 14 |
TP_SIZE="${TP_SIZE:-1}" # this model fits 1 GPU β keep 1
|
|
|
|
| 8 |
|
| 9 |
# --- tunables (all env-overridable) ---
|
| 10 |
MAX_MODEL_LEN="${MAX_MODEL_LEN:-32768}" # up to 128000
|
| 11 |
+
MAX_NUM_SEQS="${MAX_NUM_SEQS:-16}" # ~10 full-128KθΆ³θ»½ + headroom for shorter ones (fp8 KV)
|
| 12 |
GPU_MEM_UTIL="${GPU_MEM_UTIL:-0.90}"
|
| 13 |
KV_CACHE_DTYPE="${KV_CACHE_DTYPE:-auto}" # set "fp8" to ~double KV capacity
|
| 14 |
TP_SIZE="${TP_SIZE:-1}" # this model fits 1 GPU β keep 1
|