sakamakismile commited on
Commit
48876b9
Β·
verified Β·
1 Parent(s): b9588eb

Default max-num-seqs=16, fp8-KV 128K serve example, measured concurrency (~10 full-128K sessions)

Browse files
Files changed (4) hide show
  1. README.md +8 -2
  2. USAGE.md +4 -3
  3. compose.yaml +1 -1
  4. entrypoint.sh +1 -1
README.md CHANGED
@@ -86,13 +86,19 @@ answer in `content`. Omit it if you want the raw text (think tags included).
86
  CUDA_VISIBLE_DEVICES=0 vllm serve sakamakismile/LFM2.5-8B-A1B-NVFP4 \
87
  --served-model-name lfm25-8b-a1b \
88
  --quantization modelopt \
89
- --max-model-len 32768 \
90
- --max-num-seqs 8 \
 
91
  --gpu-memory-utilization 0.90 \
92
  --reasoning-parser deepseek_r1 \
93
  --port 8000
94
  ```
95
 
 
 
 
 
 
96
  | flag | what it does for *this* model |
97
  |---|---|
98
  | `--quantization modelopt` | **required** β€” reads `hf_quant_config.json` (NVFP4). Omit it and weights load as raw uint8 β†’ garbage. |
 
86
  CUDA_VISIBLE_DEVICES=0 vllm serve sakamakismile/LFM2.5-8B-A1B-NVFP4 \
87
  --served-model-name lfm25-8b-a1b \
88
  --quantization modelopt \
89
+ --kv-cache-dtype fp8 \
90
+ --max-model-len 128000 \
91
+ --max-num-seqs 16 \
92
  --gpu-memory-utilization 0.90 \
93
  --reasoning-parser deepseek_r1 \
94
  --port 8000
95
  ```
96
 
97
+ **Concurrency on one 16 GB card (measured):** with `--kv-cache-dtype fp8` the KV pool is
98
+ **1.38 M tokens** β†’ **~10 sessions each at a full 128 K context** (bf16 KV β†’ ~5). With
99
+ paged KV and shorter prompts you can serve far more β€” `--max-num-seqs 16` is a good
100
+ default; pin to `10` to guarantee every slot at full 128 K.
101
+
102
  | flag | what it does for *this* model |
103
  |---|---|
104
  | `--quantization modelopt` | **required** β€” reads `hf_quant_config.json` (NVFP4). Omit it and weights load as raw uint8 β†’ garbage. |
USAGE.md CHANGED
@@ -11,8 +11,9 @@ Tested: vLLM 0.21, CUDA 12.8/13.x, SM120. ~117 tok/s single, ~326 tok/s @4 concu
11
  CUDA_VISIBLE_DEVICES=0 vllm serve /path/to/LFM2.5-8B-A1B-NVFP4 \
12
  --served-model-name lfm25-8b-a1b \
13
  --quantization modelopt \
14
- --max-model-len 32768 \
15
- --max-num-seqs 8 \
 
16
  --gpu-memory-utilization 0.90 \
17
  --reasoning-parser deepseek_r1 \
18
  --port 8000
@@ -34,7 +35,7 @@ curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -
34
  |---|---|---|
35
  | `--quantization modelopt` | **required** | tells vLLM the checkpoint is NVFP4 (`hf_quant_config.json`). Without it vLLM treats weights as raw uint8 β†’ garbage. |
36
  | `--max-model-len` | `32768` (up to `128000`) | context window. KV is cheap here (only 6 of 24 layers carry KV), so you can go long or run many sessions. |
37
- | `--max-num-seqs` | `8` (try `16`+) | concurrent sequences. Small weights (6.9 GB) leave ~7.5 GB for KV β‰ˆ **658 K tokens**, so this card holds *a lot* of parallel sessions β€” raise it for throughput. |
38
  | `--gpu-memory-utilization` | `0.90` | fraction of VRAM vLLM may use. 0.90 is safe on a clean 16 GB card; lower if the card is shared. |
39
  | `--reasoning-parser deepseek_r1` | recommended | LFM2.5 is a **reasoning model** emitting `<think>…</think>` then the answer. This parser puts the CoT in `reasoning_content` and the answer in `content`. Omit it to get the raw text (think tags included). |
40
  | `--kv-cache-dtype fp8` | optional | ~doubles KV capacity (β†’ ~1.3 M tokens) for even more concurrency; tiny quality cost. Default (omit) keeps bf16 KV. |
 
11
  CUDA_VISIBLE_DEVICES=0 vllm serve /path/to/LFM2.5-8B-A1B-NVFP4 \
12
  --served-model-name lfm25-8b-a1b \
13
  --quantization modelopt \
14
+ --kv-cache-dtype fp8 \
15
+ --max-model-len 128000 \
16
+ --max-num-seqs 16 \
17
  --gpu-memory-utilization 0.90 \
18
  --reasoning-parser deepseek_r1 \
19
  --port 8000
 
35
  |---|---|---|
36
  | `--quantization modelopt` | **required** | tells vLLM the checkpoint is NVFP4 (`hf_quant_config.json`). Without it vLLM treats weights as raw uint8 β†’ garbage. |
37
  | `--max-model-len` | `32768` (up to `128000`) | context window. KV is cheap here (only 6 of 24 layers carry KV), so you can go long or run many sessions. |
38
+ | `--max-num-seqs` | `16` | concurrent sequences. Measured on a 16 GB card: with `--kv-cache-dtype fp8` the KV pool is **1.38 M tokens** β†’ **~10 sessions each at a FULL 128 K context** (bf16 KV β†’ ~5). 16 leaves headroom for shorter sessions via paged KV; pin to `10` if you want every slot guaranteed at full 128 K. |
39
  | `--gpu-memory-utilization` | `0.90` | fraction of VRAM vLLM may use. 0.90 is safe on a clean 16 GB card; lower if the card is shared. |
40
  | `--reasoning-parser deepseek_r1` | recommended | LFM2.5 is a **reasoning model** emitting `<think>…</think>` then the answer. This parser puts the CoT in `reasoning_content` and the answer in `content`. Omit it to get the raw text (think tags included). |
41
  | `--kv-cache-dtype fp8` | optional | ~doubles KV capacity (β†’ ~1.3 M tokens) for even more concurrency; tiny quality cost. Default (omit) keeps bf16 KV. |
compose.yaml CHANGED
@@ -13,7 +13,7 @@ services:
13
  - MODEL_DIR=/model
14
  - PORT=8000
15
  - MAX_MODEL_LEN=${MAX_MODEL_LEN:-32768}
16
- - MAX_NUM_SEQS=${MAX_NUM_SEQS:-8}
17
  - GPU_MEM_UTIL=${GPU_MEM_UTIL:-0.90}
18
  - KV_CACHE_DTYPE=${KV_CACHE_DTYPE:-auto}
19
  - TP_SIZE=${TP_SIZE:-1}
 
13
  - MODEL_DIR=/model
14
  - PORT=8000
15
  - MAX_MODEL_LEN=${MAX_MODEL_LEN:-32768}
16
+ - MAX_NUM_SEQS=${MAX_NUM_SEQS:-16}
17
  - GPU_MEM_UTIL=${GPU_MEM_UTIL:-0.90}
18
  - KV_CACHE_DTYPE=${KV_CACHE_DTYPE:-auto}
19
  - TP_SIZE=${TP_SIZE:-1}
entrypoint.sh CHANGED
@@ -8,7 +8,7 @@ HOST="${HOST:-0.0.0.0}"
8
 
9
  # --- tunables (all env-overridable) ---
10
  MAX_MODEL_LEN="${MAX_MODEL_LEN:-32768}" # up to 128000
11
- MAX_NUM_SEQS="${MAX_NUM_SEQS:-8}" # raise for more concurrency (KV is cheap here)
12
  GPU_MEM_UTIL="${GPU_MEM_UTIL:-0.90}"
13
  KV_CACHE_DTYPE="${KV_CACHE_DTYPE:-auto}" # set "fp8" to ~double KV capacity
14
  TP_SIZE="${TP_SIZE:-1}" # this model fits 1 GPU β€” keep 1
 
8
 
9
  # --- tunables (all env-overridable) ---
10
  MAX_MODEL_LEN="${MAX_MODEL_LEN:-32768}" # up to 128000
11
+ MAX_NUM_SEQS="${MAX_NUM_SEQS:-16}" # ~10 full-128KθΆ³θ»½ + headroom for shorter ones (fp8 KV)
12
  GPU_MEM_UTIL="${GPU_MEM_UTIL:-0.90}"
13
  KV_CACHE_DTYPE="${KV_CACHE_DTYPE:-auto}" # set "fp8" to ~double KV capacity
14
  TP_SIZE="${TP_SIZE:-1}" # this model fits 1 GPU β€” keep 1