djdeniro commited on
Commit
7bfe6c8
·
verified ·
1 Parent(s): 9c3c6bd

Remove duplicate vllm_deploy_guide.md from root (exists in docs/)

Browse files
Files changed (1) hide show
  1. vllm_deploy_guide.md +0 -193
vllm_deploy_guide.md DELETED
@@ -1,193 +0,0 @@
1
- # MiniMax-M2.7-MXFP416 vLLM Deployment Guide (RDNA 4 / RX 9700)
2
-
3
- ## Hardware Requirements
4
-
5
- - **GPU:** AMD RX 9700 (RDNA 4 / gfx12xx) — minimum 4x recommended
6
- - **Memory:** 128GB+ system RAM for 8-GPU setup
7
- - **OS:** Linux with ROCm 6.x
8
- - **Docker:** RDNA4-compatible vLLM image
9
-
10
- ## Docker Image
11
-
12
- The only validated runtime for this model is [`tcclaviger/vllm22:latest`](https://hub.docker.com/r/tcclaviger/vllm22):
13
-
14
- ```bash
15
- docker pull tcclaviger/vllm22:latest
16
- ```
17
-
18
- This image includes:
19
- - Custom Triton attention kernels tuned for RDNA4 (significantly faster than ROCm attention at long context)
20
- - Fixed FP8 KV-cache quantization path (~2× throughput improvement)
21
- - Pre-tuned GEMM configs for RX 9700
22
- - MXFP4-16 kernels compiled for gfx12xx
23
-
24
- ## System Setup
25
-
26
- ### GPU Devices
27
-
28
- Make sure all GPUs are visible:
29
-
30
- ```bash
31
- rocm-smi --showid
32
- # Should show: 0, 1, 2, 3, 4, 5, 6, 7
33
- ```
34
-
35
- ### Power Limit (Recommended)
36
-
37
- RDNA4 performs best with tuned power limits. Default is ~300W but 210W provides better sustained throughput on multi-GPU setups:
38
-
39
- ```bash
40
- # Set per-GPU power limit
41
- for i in 0 1 2 3 4 5 6 7; do
42
- rocm-smi --setpowerlimit $i 210
43
- done
44
- ```
45
-
46
- > **Note:** At full power (300W) sustained speeds are lower due to thermal throttling. At 210W, sustained generation throughput is consistently higher under multi-user workloads.
47
-
48
- ## Launching the Server
49
-
50
- ### Single Container (8 GPUs)
51
-
52
- ```bash
53
- docker run --name minimax-mxfp416 \
54
- --rm --tty --ipc=host --shm-size=128g \
55
- --device /dev/kfd:/dev/kfd \
56
- --device /dev/dri/renderD128:/dev/dri/renderD128 \
57
- --device /dev/dri/renderD129:/dev/dri/renderD129 \
58
- --device /dev/dri/renderD130:/dev/dri/renderD130 \
59
- --device /dev/dri/renderD132:/dev/dri/renderD132 \
60
- --device /dev/dri/renderD137:/dev/dri/renderD137 \
61
- --device /dev/dri/renderD138:/dev/dri/renderD138 \
62
- --device /dev/dri/renderD139:/dev/dri/renderD139 \
63
- --device /dev/dri/renderD140:/dev/dri/renderD140 \
64
- -e HIP_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
65
- -e ROCR_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
66
- -e TRUST_REMOTE_CODE=1 \
67
- -e PYTORCH_TUNABLEOP_ENABLED=1 \
68
- -e PYTORCH_TUNABLEOP_TUNING=0 \
69
- -e PYTORCH_TUNABLEOP_RECORD_UNTUNED=0 \
70
- -e PYTORCH_ALLOC_CONF=expandable_segments:True \
71
- -e PYTORCH_HIP_ALLOC_CONF=expandable_segments:True \
72
- -e GPU_MAX_HW_QUEUES=1 \
73
- -v /path/to/models:/app/models:ro \
74
- -v /path/to/patches:/patches:ro \
75
- -p 8000:8000 \
76
- tcclaviger/vllm22:latest \
77
- bash -c "cp /patches/vllm22_minimax_m2.py /app/vllm/vllm/model_executor/models/minimax_m2.py && \
78
- pip install -q sentencepiece && \
79
- exec vllm serve \
80
- /app/models/MiniMax-M2.7-MXFP416 \
81
- --served-model-name minimax-m2.7-mxfp416 \
82
- --host 0.0.0.0 --port 8000 \
83
- --trust-remote-code \
84
- --tensor-parallel-size 8 \
85
- --enable-expert-parallel \
86
- --disable-cascade-attn \
87
- --reasoning-parser minimax_m2 \
88
- --enable-auto-tool-choice \
89
- --tool-call-parser minimax_m2 \
90
- --enable-prefix-caching \
91
- --gpu-memory-utilization 0.93 \
92
- --max-model-len 180000 \
93
- --max-num-seqs 48 \
94
- --max-num-batched-tokens 2048 \
95
- --kv-cache-dtype fp8_e4m3 \
96
- --attention-backend TRITON_ATTN \
97
- --override-generation-config '{\"max_tokens\": 16384}'"
98
- ```
99
-
100
- ### Key vLLM Flags Explained
101
-
102
- | Flag | Value | Purpose |
103
- |------|-------|---------|
104
- | `--tensor-parallel-size` | 8 | Split model across 8 GPUs |
105
- | `--enable-expert-parallel` | | Enable expert-parallel distribution |
106
- | `--disable-cascade-attn` | | Disable cascade attention for MoE layers |
107
- | `--attention-backend` | TRITON_ATTN | Use Triton kernels (10× faster than ROCm on RDNA4) |
108
- | `--kv-cache-dtype` | fp8_e4m3 | FP8 KV cache (~50% memory savings) |
109
- | `--enable-prefix-caching` | | Cache common prefixes (93%+ hit rate observed) |
110
- | `--max-model-len` | 180000 | 180k context |
111
- | `--max-num-seqs` | 48 | Max concurrent sequences |
112
- | `--max-num-batched-tokens` | 2048 | Max tokens per batch |
113
- | `--gpu-memory-utilization` | 0.93 | Use 93% of GPU memory |
114
-
115
- > **Important:** The `--disable-cascade-attn` flag is required for MoE models. Without it, the model will produce incorrect outputs.
116
-
117
- ### Running with Patches
118
-
119
- If you have custom model patches:
120
-
121
- ```bash
122
- -v /path/to/patches:/patches:ro \
123
- ```
124
-
125
- The Docker entry point copies `vllm22_minimax_m2.py` to the vLLM model directory before launching. This adds MXFP4-16 support for MiniMax-M2.7.
126
-
127
- ## Performance Notes
128
-
129
- ### Observed Performance (4× RX 9700, 210W power limit)
130
-
131
- - **Generation throughput:** 50–80 tokens/s
132
- - **Prefill throughput:** 2000+ tokens/s (with prefix caching)
133
- - **Prefix cache hit rate:** ~93%
134
- - **KV cache usage:** 25–33% typical at 180k context
135
- - **Max concurrent users:** 4–5 at full 180k context
136
-
137
- ### KV Cache Capacity
138
-
139
- With 8× RX 9700 and FP8 KV cache:
140
- - **KV cache memory:** 11.35 GiB
141
- - **KV cache tokens:** ~768K tokens
142
- - **Max context per request:** 180,000 tokens
143
- - **Max concurrent at 180k:** ~4 requests
144
-
145
- ### Model Loading
146
-
147
- - **Weight loading time:** ~42 seconds
148
- - **Memory per GPU (TP8):** ~17.5 GiB
149
- - **Torch compile warmup:** ~37 seconds
150
-
151
- ## Testing the Deployment
152
-
153
- ```python
154
- from openai import OpenAI
155
-
156
- client = OpenAI(
157
- base_url="http://localhost:8000/v1",
158
- api_key="EMPTY",
159
- )
160
-
161
- completion = client.chat.completions.create(
162
- model="minimax-m2.7-mxfp416",
163
- messages=[
164
- {"role": "system", "content": "You are a helpful assistant."},
165
- {"role": "user", "content": "Explain what MXFP4 quantization is in one sentence."}
166
- ],
167
- temperature=1.0,
168
- max_tokens=256,
169
- )
170
-
171
- print(completion.choices[0].message.content)
172
- ```
173
-
174
- ## Troubleshooting
175
-
176
- ### "expandable_segments not supported"
177
-
178
- This warning is benign on ROCm. The model runs correctly despite the warning.
179
-
180
- ### Low throughput at long context
181
-
182
- Ensure `TRITON_ATTN` backend is active. Default ROCm attention is 10× slower on RDNA4 at long context.
183
-
184
- ### Thermal throttling
185
-
186
- If sustained throughput degrades over time, reduce power limit to 210W per GPU:
187
- ```bash
188
- rocm-smi --setpowerlimit 0 210
189
- ```
190
-
191
- ### Model fails to load
192
-
193
- Ensure `--trust-remote-code` is set and the model path is correct. The custom model file (`vllm22_minimax_m2.py`) must be copied before vLLM loads the model.