Capicua25x commited on
Commit
dd8bb26
Β·
verified Β·
1 Parent(s): 532de84

remove single-card section pending further evaluation

Browse files
Files changed (1) hide show
  1. README.md +0 -47
README.md CHANGED
@@ -163,53 +163,6 @@ not drafting. What C buys instead: the full 262k native window (FP8+MTP tops out
163
  Pick FP8 for short-query burst capacity; pick this model for max context on the same silicon β€”
164
  at real context sizes you give up nothing measurable.
165
 
166
- ## Single card (1Γ— R9700 / 32GB-class)
167
-
168
- **Single-card path: MXFP4 @ fp8 KV** β€” the only quant that fits (weights β‰ˆ21 GB); fp8 KV stretches
169
- the window from 32k to 104k. It serves β€” with real limits: β‰ˆ4 concurrent users and a window still
170
- well short of the 262k the multi-card port gets on the same silicon. llama.cpp/GGUF can also run
171
- this model on one card for plain chat, but skips the batching, prefix caching, and window this
172
- config gets from the same card β€” not the recommendation unless that's genuinely all you need.
173
-
174
- Measured 2026-08-25 on 1Γ— R9700, rc10 image, bench v4, KV pool 113,642 tokens (127,348 at 4 slots)
175
- β†’ 104k-token window:
176
-
177
- | workload (v4, TP1 + fp8 KV, 104k window, 4 slots) | c1 | c2 | c4 |
178
- |---|---|---|---|
179
- | trivial | 43.8 / 44 | 40.7 / 81 | 39.9 / 143 |
180
- | short essay | 29.0 / 29 | 27.6 / 51 | 23.5 / 85 |
181
- | 6k-prefix essay | 32.4 / 32 | 27.0 / 51 | 23.7 / 90 |
182
-
183
- Stays β‰₯20 tok/s/user through c4 (untested beyond). Serve command β€” one device pair, TP1, 104k
184
- window, 4 slots:
185
-
186
- ```bash
187
- docker run --rm --name vllm-qwen --network=host \
188
- --device=/dev/kfd --device=/dev/dri/renderD128 --device=/dev/dri/card1 \
189
- --group-add=video --group-add=render --ipc=host \
190
- -v /path/to/quants:/quant:ro -e HF_HUB_OFFLINE=1 \
191
- --entrypoint /usr/local/bin/vllm capicua25x/vllm-rocm-rdna4:latest \
192
- serve /quant/Qwen3.8-27B-MXFP4-Quark-RDNA4 \
193
- --served-model-name qwen --port 8011 --trust-remote-code \
194
- --tensor-parallel-size 1 --gpu-memory-utilization 0.95 --max-model-len 106496 \
195
- --attention-backend TRITON_ATTN --enable-prefix-caching \
196
- --max-num-seqs 4 --max-num-batched-tokens 8000 \
197
- --kv-cache-dtype fp8 --mamba-ssm-cache-dtype bfloat16 --max-cudagraph-capture-size 32 \
198
- --speculative-config '{"method":"mtp","num_speculative_tokens":3,"attention_backend":"TRITON_ATTN"}'
199
- ```
200
-
201
- **16GB cards (RX 9070 XT): this 27B does not fit β€” in ANY runtime.** Our quants need 21 GB of
202
- weights; even a Q4 GGUF (β‰ˆ15.5 GB) technically loads and then leaves no room for KV, i.e. no
203
- usable context window. llama.cpp's 16GB options are CPU offload (at a large speed cost) or a
204
- smaller model β€” the latter is the honest recommendation. The port's kernels run fine on gfx1200
205
- silicon; pair the card with a model whose weights + context leave real headroom.
206
-
207
- Quality note: fp8 KV over this MXFP4 quant carries one accuracy smoke (gsm8k n=50 thinking-on:
208
- 0.94 flexible / 0.88 strict β€” within the Β±0.04 sampling noise of the bf16-KV baseline's 0.91–0.93,
209
- at its lower edge). Validate on your own workload before committing; a KV-calibrated variant with
210
- scale side-files exists if deeper validation matters to you.
211
-
212
-
213
  ## Accuracy (AA class-A, paired items, seed 1234, on-spec sampling)
214
 
215
  **ref** = the same checkpoint served in bf16 by a cloud provider. Same judge for all judged rows.
 
163
  Pick FP8 for short-query burst capacity; pick this model for max context on the same silicon β€”
164
  at real context sizes you give up nothing measurable.
165
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
166
  ## Accuracy (AA class-A, paired items, seed 1234, on-spec sampling)
167
 
168
  **ref** = the same checkpoint served in bf16 by a cloud provider. Same judge for all judged rows.