djdeniro commited on
Commit
07cb9b1
·
verified ·
1 Parent(s): 4f82b05

Add RDNA4/RX9700 vLLM deployment guide

Browse files
Files changed (1) hide show
  1. docs/vllm_deploy_guide.md +193 -0
docs/vllm_deploy_guide.md ADDED
@@ -0,0 +1,193 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # MiniMax-M2.7-MXFP416 vLLM Deployment Guide (RDNA 4 / RX 9700)
2
+
3
+ ## Hardware Requirements
4
+
5
+ - **GPU:** AMD RX 9700 (RDNA 4 / gfx12xx) — minimum 4x recommended
6
+ - **Memory:** 128GB+ system RAM for 8-GPU setup
7
+ - **OS:** Linux with ROCm 6.x
8
+ - **Docker:** RDNA4-compatible vLLM image
9
+
10
+ ## Docker Image
11
+
12
+ The only validated runtime for this model is [`tcclaviger/vllm22:latest`](https://hub.docker.com/r/tcclaviger/vllm22):
13
+
14
+ ```bash
15
+ docker pull tcclaviger/vllm22:latest
16
+ ```
17
+
18
+ This image includes:
19
+ - Custom Triton attention kernels tuned for RDNA4 (significantly faster than ROCm attention at long context)
20
+ - Fixed FP8 KV-cache quantization path (~2× throughput improvement)
21
+ - Pre-tuned GEMM configs for RX 9700
22
+ - MXFP4-16 kernels compiled for gfx12xx
23
+
24
+ ## System Setup
25
+
26
+ ### GPU Devices
27
+
28
+ Make sure all GPUs are visible:
29
+
30
+ ```bash
31
+ rocm-smi --showid
32
+ # Should show: 0, 1, 2, 3, 4, 5, 6, 7
33
+ ```
34
+
35
+ ### Power Limit (Recommended)
36
+
37
+ RDNA4 performs best with tuned power limits. Default is ~300W but 210W provides better sustained throughput on multi-GPU setups:
38
+
39
+ ```bash
40
+ # Set per-GPU power limit
41
+ for i in 0 1 2 3 4 5 6 7; do
42
+ rocm-smi --setpowerlimit $i 210
43
+ done
44
+ ```
45
+
46
+ > **Note:** At full power (300W) sustained speeds are lower due to thermal throttling. At 210W, sustained generation throughput is consistently higher under multi-user workloads.
47
+
48
+ ## Launching the Server
49
+
50
+ ### Single Container (8 GPUs)
51
+
52
+ ```bash
53
+ docker run --name minimax-mxfp416 \
54
+ --rm --tty --ipc=host --shm-size=128g \
55
+ --device /dev/kfd:/dev/kfd \
56
+ --device /dev/dri/renderD128:/dev/dri/renderD128 \
57
+ --device /dev/dri/renderD129:/dev/dri/renderD129 \
58
+ --device /dev/dri/renderD130:/dev/dri/renderD130 \
59
+ --device /dev/dri/renderD132:/dev/dri/renderD132 \
60
+ --device /dev/dri/renderD137:/dev/dri/renderD137 \
61
+ --device /dev/dri/renderD138:/dev/dri/renderD138 \
62
+ --device /dev/dri/renderD139:/dev/dri/renderD139 \
63
+ --device /dev/dri/renderD140:/dev/dri/renderD140 \
64
+ -e HIP_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
65
+ -e ROCR_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
66
+ -e TRUST_REMOTE_CODE=1 \
67
+ -e PYTORCH_TUNABLEOP_ENABLED=1 \
68
+ -e PYTORCH_TUNABLEOP_TUNING=0 \
69
+ -e PYTORCH_TUNABLEOP_RECORD_UNTUNED=0 \
70
+ -e PYTORCH_ALLOC_CONF=expandable_segments:True \
71
+ -e PYTORCH_HIP_ALLOC_CONF=expandable_segments:True \
72
+ -e GPU_MAX_HW_QUEUES=1 \
73
+ -v /path/to/models:/app/models:ro \
74
+ -v /path/to/patches:/patches:ro \
75
+ -p 8000:8000 \
76
+ tcclaviger/vllm22:latest \
77
+ bash -c "cp /patches/vllm22_minimax_m2.py /app/vllm/vllm/model_executor/models/minimax_m2.py && \
78
+ pip install -q sentencepiece && \
79
+ exec vllm serve \
80
+ /app/models/MiniMax-M2.7-MXFP416 \
81
+ --served-model-name minimax-m2.7-mxfp416 \
82
+ --host 0.0.0.0 --port 8000 \
83
+ --trust-remote-code \
84
+ --tensor-parallel-size 8 \
85
+ --enable-expert-parallel \
86
+ --disable-cascade-attn \
87
+ --reasoning-parser minimax_m2 \
88
+ --enable-auto-tool-choice \
89
+ --tool-call-parser minimax_m2 \
90
+ --enable-prefix-caching \
91
+ --gpu-memory-utilization 0.93 \
92
+ --max-model-len 180000 \
93
+ --max-num-seqs 48 \
94
+ --max-num-batched-tokens 2048 \
95
+ --kv-cache-dtype fp8_e4m3 \
96
+ --attention-backend TRITON_ATTN \
97
+ --override-generation-config '{\"max_tokens\": 16384}'"
98
+ ```
99
+
100
+ ### Key vLLM Flags Explained
101
+
102
+ | Flag | Value | Purpose |
103
+ |------|-------|---------|
104
+ | `--tensor-parallel-size` | 8 | Split model across 8 GPUs |
105
+ | `--enable-expert-parallel` | | Enable expert-parallel distribution |
106
+ | `--disable-cascade-attn` | | Disable cascade attention for MoE layers |
107
+ | `--attention-backend` | TRITON_ATTN | Use Triton kernels (10× faster than ROCm on RDNA4) |
108
+ | `--kv-cache-dtype` | fp8_e4m3 | FP8 KV cache (~50% memory savings) |
109
+ | `--enable-prefix-caching` | | Cache common prefixes (93%+ hit rate observed) |
110
+ | `--max-model-len` | 180000 | 180k context |
111
+ | `--max-num-seqs` | 48 | Max concurrent sequences |
112
+ | `--max-num-batched-tokens` | 2048 | Max tokens per batch |
113
+ | `--gpu-memory-utilization` | 0.93 | Use 93% of GPU memory |
114
+
115
+ > **Important:** The `--disable-cascade-attn` flag is required for MoE models. Without it, the model will produce incorrect outputs.
116
+
117
+ ### Running with Patches
118
+
119
+ If you have custom model patches:
120
+
121
+ ```bash
122
+ -v /path/to/patches:/patches:ro \
123
+ ```
124
+
125
+ The Docker entry point copies `vllm22_minimax_m2.py` to the vLLM model directory before launching. This adds MXFP4-16 support for MiniMax-M2.7.
126
+
127
+ ## Performance Notes
128
+
129
+ ### Observed Performance (4× RX 9700, 210W power limit)
130
+
131
+ - **Generation throughput:** 50–80 tokens/s
132
+ - **Prefill throughput:** 2000+ tokens/s (with prefix caching)
133
+ - **Prefix cache hit rate:** ~93%
134
+ - **KV cache usage:** 25–33% typical at 180k context
135
+ - **Max concurrent users:** 4–5 at full 180k context
136
+
137
+ ### KV Cache Capacity
138
+
139
+ With 8× RX 9700 and FP8 KV cache:
140
+ - **KV cache memory:** 11.35 GiB
141
+ - **KV cache tokens:** ~768K tokens
142
+ - **Max context per request:** 180,000 tokens
143
+ - **Max concurrent at 180k:** ~4 requests
144
+
145
+ ### Model Loading
146
+
147
+ - **Weight loading time:** ~42 seconds
148
+ - **Memory per GPU (TP8):** ~17.5 GiB
149
+ - **Torch compile warmup:** ~37 seconds
150
+
151
+ ## Testing the Deployment
152
+
153
+ ```python
154
+ from openai import OpenAI
155
+
156
+ client = OpenAI(
157
+ base_url="http://localhost:8000/v1",
158
+ api_key="EMPTY",
159
+ )
160
+
161
+ completion = client.chat.completions.create(
162
+ model="minimax-m2.7-mxfp416",
163
+ messages=[
164
+ {"role": "system", "content": "You are a helpful assistant."},
165
+ {"role": "user", "content": "Explain what MXFP4 quantization is in one sentence."}
166
+ ],
167
+ temperature=1.0,
168
+ max_tokens=256,
169
+ )
170
+
171
+ print(completion.choices[0].message.content)
172
+ ```
173
+
174
+ ## Troubleshooting
175
+
176
+ ### "expandable_segments not supported"
177
+
178
+ This warning is benign on ROCm. The model runs correctly despite the warning.
179
+
180
+ ### Low throughput at long context
181
+
182
+ Ensure `TRITON_ATTN` backend is active. Default ROCm attention is 10× slower on RDNA4 at long context.
183
+
184
+ ### Thermal throttling
185
+
186
+ If sustained throughput degrades over time, reduce power limit to 210W per GPU:
187
+ ```bash
188
+ rocm-smi --setpowerlimit 0 210
189
+ ```
190
+
191
+ ### Model fails to load
192
+
193
+ Ensure `--trust-remote-code` is set and the model path is correct. The custom model file (`vllm22_minimax_m2.py`) must be copied before vLLM loads the model.