philbert440 commited on
Commit
c36af8e
·
verified ·
1 Parent(s): df36d26

Model card: V100 measured performance (warm multi-pass), draft-mode + cudagraph sizing guidance, chart

Browse files
Files changed (2) hide show
  1. README.md +34 -4
  2. images/throughput-v100.png +0 -0
README.md CHANGED
@@ -53,9 +53,38 @@ ModelOpt-exported NVFP4 checkpoints require capability 7.5+ and reject Volta. On
53
 
54
  ## Measured performance
55
 
56
- Benchmarks on 2×V100-32GB (1Cat-vLLM 1.2.2, TP2, `fp8_e5m2` KV, official Qwen3.8 thinking-mode
57
- sampling) are being run now — multi-pass, warm-serve methodology, MTP K=2 in both draft modes.
58
- This section will be updated with the results and a throughput chart.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
59
 
60
  ## The base model
61
 
@@ -122,7 +151,8 @@ python -m vllm.entrypoints.openai.api_server \
122
  --tensor-parallel-size 2 --gpu-memory-utilization 0.78 \
123
  --max-model-len 32768 --kv-cache-dtype fp8_e5m2 \
124
  --enable-prefix-caching --reasoning-parser qwen3 \
125
- --speculative-config '{"method":"mtp","num_speculative_tokens":2,"attention_backend":"FLASH_ATTN_V100","draft_sample_method":"greedy"}'
 
126
  ```
127
 
128
  SM70 notes, learned the hard way:
 
53
 
54
  ## Measured performance
55
 
56
+ ![MTP draft-mode throughput](images/throughput-v100.png)
57
+
58
+ **Methodology:** warm serve (2 discarded warmup generations), fixed-length generations via
59
+ `ignore_eos` so every run produces exactly the stated token count, official Qwen3.8 sampling per
60
+ mode (thinking `1.0/0.95/20`, instruct `0.7/0.80/20 + presence 1.5`), varied prompts. Reported
61
+ as mean ± sd tokens/s. Rig: 2×V100-32GB, 1Cat-vLLM 1.2.2, TP2, `fp8_e5m2` KV,
62
+ `max_num_seqs 4`, MTP K=2.
63
+
64
+ | Regime | greedy draft | probabilistic draft |
65
+ |---|---|---|
66
+ | 512-tok, thinking (n=10) | 53.0 ± 2.5 | **54.4 ± 1.3** |
67
+ | 2048-tok, thinking (n=3) | 51.1 ± 2.5 | **54.1 ± 1.1** |
68
+ | 512-tok, instruct (n=6) | 50.5 ± 1.8 | 50.0 ± 1.6 |
69
+ | Mean acceptance length, whole workload | 2.25 | 2.46 |
70
+
71
+ **Concurrency** (4-way, 512-tok, aggregate): **~170 tok/s** with
72
+ `{"cudagraph_mode":"piecewise"}` (auto capture sizes), ~165 with `full_and_piecewise` — on par
73
+ with the [W4A16 sibling](https://huggingface.co/philbert440/Qwen3.8-27B-W4A16-AWQ). One
74
+ sizing rule matters: with MTP, each sequence schedules `K+1` tokens per step, so **never set
75
+ explicit `cudagraph_capture_sizes` below `max_num_seqs × (K+1)`** — a cap of `[1,2,4,8]` at
76
+ batch 4 pushes concurrent decode off CUDA graphs and collapses aggregate throughput ~3×
77
+ (measured 55–72 tok/s; reproduces identically on the W4A16 sibling, so it's a config trap, not
78
+ a format property). Engine-default auto sizing is correct.
79
+
80
+ **Pick the draft mode by workload:** verification rejection-samples against the target model,
81
+ so output quality is identical either way. At the official temp-1.0 thinking sampling,
82
+ **probabilistic** matches the verified distribution and wins (+2–6%); on low-temperature
83
+ workloads the two converge (see instruct row); at `temperature 0` greedy is the natural choice.
84
+
85
+ Quality validation (passed on this rig): factual coherence, think-tag discipline (zero
86
+ `<think>` leakage with thinking disabled), vision (image understanding through the VLM path),
87
+ GSM8K sample 3/3, and long-form generation with no repetition/degeneration.
88
 
89
  ## The base model
90
 
 
151
  --tensor-parallel-size 2 --gpu-memory-utilization 0.78 \
152
  --max-model-len 32768 --kv-cache-dtype fp8_e5m2 \
153
  --enable-prefix-caching --reasoning-parser qwen3 \
154
+ --compilation-config '{"cudagraph_mode":"piecewise"}' \
155
+ --speculative-config '{"method":"mtp","num_speculative_tokens":2,"attention_backend":"FLASH_ATTN_V100","draft_sample_method":"probabilistic"}'
156
  ```
157
 
158
  SM70 notes, learned the hard way:
images/throughput-v100.png ADDED