djdeniro commited on
Commit
7c20f79
·
verified ·
1 Parent(s): 1dc5d39

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +26 -7
README.md CHANGED
@@ -60,6 +60,18 @@ served with **vLLM** on **RDNA4** (AMD Radeon R9700) hardware.
60
 
61
  ***Total Context Limit for each task in test 32k, means 6x tasks use more than 32k output tokens***
62
 
 
 
 
 
 
 
 
 
 
 
 
 
63
 
64
  ---
65
 
@@ -125,15 +137,20 @@ docker pull tcclaviger/vllm:latest
125
 
126
  git clone https://huggingface.co/djdeniro/GLM-5.3-Flash-rocm-r9700 overlay
127
 
 
128
  docker run --rm --tty --ipc=host --shm-size=128g \
129
  --device /dev/kfd:/dev/kfd --device /dev/dri:/dev/dri \
130
  -v /path/to/GLM-5.3-Flash-RFA-RFI8-8xR9700:/models:ro \
131
  -v "$PWD/overlay":/overlay:ro \
132
  --entrypoint bash tcclaviger/vllm:latest \
133
- -c "/overlay/apply_overlay.sh && exec vllm serve /models \
134
  --served-model-name glm53-flash --trust-remote-code --quantization rfi \
135
- --tensor-parallel-size 8 --gpu-memory-utilization 0.95 \
136
- --max-model-len 190080 --max-num-seqs 4 --kv-cache-dtype auto"
 
 
 
 
137
  ```
138
 
139
  ---
@@ -146,7 +163,7 @@ docker run --rm --tty --ipc=host --shm-size=128g \
146
  | Layers | 45 = 34 KDA (linear attention) + 11 DSA (sparse-MLA) |
147
  | Routed experts | 288 (top-8) + 1 shared expert |
148
  | Extra | mHC hyper-connections, 1 nextn MTP draft layer, native vision tower |
149
- | Context (bf16 KV) | 190,080 tokens |
150
 
151
  ---
152
 
@@ -159,9 +176,11 @@ fed with a min/max image-token budget. The model accepts image and video inputs
159
 
160
  ## Known limitations
161
 
162
- - **MTP is disabled** in the reference serving config (drafter KV-group blocker).
163
- - **Serve with bf16 KV** (`--kv-cache-dtype auto`) — fp8 KV with runtime scale calibration is
164
- broken on this architecture (garbage scales from the uninitialized KDA recurrent state).
 
 
165
  - **Chat needs `reasoning_effort="low"`** — the default Reasoning Effort Max spends 16k+ tokens
166
  thinking before producing content on long generations.
167
 
 
60
 
61
  ***Total Context Limit for each task in test 32k, means 6x tasks use more than 32k output tokens***
62
 
63
+ ### Serving performance (8× R9700, FY2026-09 production config)
64
+
65
+ | Scenario | Throughput |
66
+ |----------|------------|
67
+ | Decode, batch size 1, MTP OFF | ~34–39 tok/s |
68
+ | Decode, batch size 1, **MTP spec=3** | **~82–88 tok/s** |
69
+ | Aggregate, 4 concurrent, MTP spec=3 | ~155 tok/s |
70
+ | Context window (fp8 KV) | 300,000 tokens |
71
+
72
+ MTP speculative decoding: mean acceptance length ~3.7–3.9 of 4 draft tokens,
73
+ average draft acceptance 91–97% (live engine metrics, GPQA-style prompts).
74
+
75
 
76
  ---
77
 
 
137
 
138
  git clone https://huggingface.co/djdeniro/GLM-5.3-Flash-rocm-r9700 overlay
139
 
140
+ # current production config (MTP spec=3, fp8 KV, 300k context)
141
  docker run --rm --tty --ipc=host --shm-size=128g \
142
  --device /dev/kfd:/dev/kfd --device /dev/dri:/dev/dri \
143
  -v /path/to/GLM-5.3-Flash-RFA-RFI8-8xR9700:/models:ro \
144
  -v "$PWD/overlay":/overlay:ro \
145
  --entrypoint bash tcclaviger/vllm:latest \
146
+ -c "/overlay/apply_overlay.sh && GLM5_NEXT_MTP_PROPOSER=1 exec vllm serve /models \
147
  --served-model-name glm53-flash --trust-remote-code --quantization rfi \
148
+ --tensor-parallel-size 8 --gpu-memory-utilization 0.9575 \
149
+ --max-model-len 300000 --max-num-seqs 4 --max-num-batched-tokens 2048 \
150
+ --kv-cache-dtype fp8 \
151
+ --speculative-config '{\"method\":\"mtp\",\"num_speculative_tokens\":3}' \
152
+ --enable-prefix-caching --distributed-executor-backend mp \
153
+ --compilation-config '{\"cudagraph_capture_sizes\":[1,2,4,8,16],\"cudagraph_mode\":\"FULL_AND_PIECEWISE\",\"cudagraph_copy_inputs\":true}'"
154
  ```
155
 
156
  ---
 
163
  | Layers | 45 = 34 KDA (linear attention) + 11 DSA (sparse-MLA) |
164
  | Routed experts | 288 (top-8) + 1 shared expert |
165
  | Extra | mHC hyper-connections, 1 nextn MTP draft layer, native vision tower |
166
+ | Context (fp8 KV) | 300,000 tokens |
167
 
168
  ---
169
 
 
176
 
177
  ## Known limitations
178
 
179
+ - **fp8 KV without runtime calibration** — serve with `--kv-cache-dtype fp8` and **scales fixed
180
+ at 1.0**. Do not enable `--calculate-kv-scales`: runtime calibration on the profile dummy-run
181
+ produces garbage scales from the uninitialized KDA recurrent state (details in the
182
+ [serving repo](https://huggingface.co/djdeniro/GLM-5.3-Flash-rocm-r9700)).
183
+ - The 300k context / MTP spec=3 config presumes the VRAM headroom of the 256 GB 8× R9700 node.
184
  - **Chat needs `reasoning_effort="low"`** — the default Reasoning Effort Max spends 16k+ tokens
185
  thinking before producing content on long generations.
186