d0xin commited on
Commit
661e557
·
verified ·
1 Parent(s): 7bf2013

Expand model card with quantization and benchmark details

Browse files
Files changed (1) hide show
  1. README.md +254 -20
README.md CHANGED
@@ -7,45 +7,279 @@ base_model: ukisai/Swift-Qwen3.8-27b
7
  base_model_relation: quantized
8
  tags:
9
  - qwen3_8
 
10
  - fp8
11
  - compressed-tensors
12
  - sglang
 
 
13
  - reasoning
 
 
14
  ---
15
 
16
  # Swift-Qwen3.8-27B-FP8
17
 
18
- FP8 quantization of `ukisai/Swift-Qwen3.8-27b`.
 
19
 
20
- Independent community quantization; not an official UkisAI release.
 
21
 
22
- ## Quantization
 
 
 
23
 
24
- - compressed-tensors
25
- - FP8_BLOCK
26
- - weight block: 128 x 128
27
- - dynamic FP8 activations
28
- - activation group size: 128
29
- - llmcompressor 0.13.0
30
- - checkpoint size: ~29 GB
31
- - context length: 262,144 tokens
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
32
 
33
- Validated with SGLang on NVIDIA RTX PRO 6000 Blackwell 96 GB.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
34
 
35
  ## Performance
36
 
37
- | Configuration | Throughput |
 
 
 
 
 
 
 
 
 
 
38
  |---|---:|
39
- | Swift FP8 | 50.45 tok/s |
40
- | Swift FP8 + NEXTN | 103.38 tok/s |
41
- | Swift FP8 + DFlash2 | **132.62 tok/s** |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
42
 
43
- DFlash2 is external and is not included in this repository.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
44
 
45
  ## License
46
 
47
- Derived from `ukisai/Swift-Qwen3.8-27b`.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
48
 
49
- This checkpoint follows the Swift Open License v1.0.
50
 
51
- https://huggingface.co/ukisai/Swift-Qwen3.8-27b
 
 
 
 
 
 
 
 
7
  base_model_relation: quantized
8
  tags:
9
  - qwen3_8
10
+ - qwen3_5
11
  - fp8
12
  - compressed-tensors
13
  - sglang
14
+ - speculative-decoding
15
+ - dflash
16
  - reasoning
17
+ - efficient-thinking
18
+ - conversational
19
  ---
20
 
21
  # Swift-Qwen3.8-27B-FP8
22
 
23
+ FP8 quantization of
24
+ [`ukisai/Swift-Qwen3.8-27b`](https://huggingface.co/ukisai/Swift-Qwen3.8-27b).
25
 
26
+ This is an independent community quantization and is **not an official
27
+ UkisAI release**.
28
 
29
+ The goal of this checkpoint is to preserve the behavior of Swift-Qwen3.8-27B
30
+ while reducing VRAM requirements and enabling high-throughput inference with
31
+ SGLang, including speculative decoding with the model's native MTP head or an
32
+ external DFlash2 draft model.
33
 
34
+ ## Model summary
35
+
36
+ - Upstream model: `ukisai/Swift-Qwen3.8-27b`
37
+ - Architecture: `Qwen3_5ForConditionalGeneration`
38
+ - Quantization format: `compressed-tensors`
39
+ - Quantization scheme: `FP8_BLOCK`
40
+ - Weight block size: `128 x 128`
41
+ - Activations: dynamic FP8
42
+ - Activation group size: `128`
43
+ - Quantizer: `llmcompressor 0.13.0`
44
+ - Declared context length: `262,144`
45
+ - Checkpoint size: approximately `29 GB`
46
+ - Native MTP components: retained
47
+
48
+ The checkpoint was produced from the BF16 Swift-Qwen3.8-27B model rather than
49
+ requantizing an already quantized derivative.
50
+
51
+ ## Quantization details
52
+
53
+ The quantization process used the official Qwen3.8 FP8 configuration as a
54
+ reference for the block-FP8 layout.
55
+
56
+ Checkpoint audit:
57
+
58
+ | Item | Count |
59
+ |---|---:|
60
+ | Total checkpoint tensors | 1,199 |
61
+ | 2D weight tensors | 617 |
62
+ | FP8 quantization candidates | 407 |
63
+ | Effectively excluded / preserved modules | 626 |
64
+ | Incompatible FP8 candidates after validation | 0 |
65
+
66
+ Matrices that are not compatible with the required `128 x 128` block
67
+ structure were preserved instead of being forcibly quantized.
68
+
69
+ The following classes of tensors were intentionally preserved where
70
+ appropriate:
71
+
72
+ - embeddings
73
+ - `lm_head`
74
+ - normalization parameters
75
+ - non-2D weights
76
+ - matrices whose dimensions are incompatible with the FP8 block layout
77
+
78
+ The native MTP layers are retained. Compatible MTP projection matrices are
79
+ quantized to FP8, while incompatible components remain unquantized.
80
+
81
+ ## Validation
82
+
83
+ Validated locally on:
84
+
85
+ - NVIDIA RTX PRO 6000 Blackwell 96 GB
86
+ - SGLang `0.5.19.dev135+ga4ffb996d`
87
+ - `compressed-tensors 0.18.0`
88
+ - CUDA-capable Linux deployment
89
+ - single-GPU tensor parallelism (`TP=1`)
90
+
91
+ SGLang successfully loads the checkpoint as:
92
 
93
+ ```text
94
+ type=Qwen3_5ForConditionalGeneration
95
+ quant=compressed-tensors
96
+ ```
97
+
98
+ Observed target-model weight memory during loading:
99
+
100
+ ```text
101
+ 28.47 GB
102
+ ```
103
+
104
+ OpenAI-compatible `/v1/chat/completions` inference was validated successfully.
105
+
106
+ Multimodal inference has not yet been separately benchmarked for this
107
+ quantized checkpoint.
108
 
109
  ## Performance
110
 
111
+ All measurements below are local measurements from a single
112
+ NVIDIA RTX PRO 6000 Blackwell 96 GB GPU.
113
+
114
+ They are intended to document this deployment, not to serve as standardized
115
+ cross-model benchmarks.
116
+
117
+ ### Fixed 4,096-token generation
118
+
119
+ Same prompt and generation setup for all configurations:
120
+
121
+ | Configuration | Median throughput |
122
  |---|---:|
123
+ | Swift FP8, target model only | 50.45 tok/s |
124
+ | Swift FP8 + native NEXTN/MTP | 103.38 tok/s |
125
+ | Swift FP8 + DFlash2, 8 draft tokens | **132.62 tok/s** |
126
+
127
+ Measured DFlash2 runs:
128
+
129
+ ```text
130
+ 131.24 tok/s
131
+ 132.65 tok/s
132
+ 132.62 tok/s
133
+ median: 132.62 tok/s
134
+ ```
135
+
136
+ Compared with target-only generation, DFlash2 produced approximately
137
+ **2.63x** higher output throughput in this test.
138
+
139
+ Compared with native NEXTN/MTP, DFlash2 was approximately **28% faster**.
140
+
141
+ The DFlash2 draft model is external and is **not included in this repository**.
142
+
143
+ ## Reasoning-heavy agent test
144
+
145
+ A separate local A/B test used the same diagnostic prompt, sampling
146
+ parameters, seed, and reasoning setting for both systems.
147
 
148
+ The prompt asked the model to diagnose an intermittently slow
149
+ OpenAI-compatible inference deployment with high GPU utilization,
150
+ large KV cache, speculative decoding, variable context sizes and
151
+ concurrency-sensitive latency.
152
+
153
+ | Metric | Previous Qwen FP8 production | Swift FP8 + DFlash2 |
154
+ |---|---:|---:|
155
+ | Wall time | 193.48 s | **139.63 s** |
156
+ | Prompt tokens | 229 | 229 |
157
+ | Reasoning tokens | 14,838 | **10,716** |
158
+ | Completion tokens | 22,848 | **16,170** |
159
+ | Finish reason | stop | stop |
160
+ | Effective completion throughput | 118.09 tok/s | 115.80 tok/s |
161
+
162
+ Observed change:
163
+
164
+ - wall-clock time: approximately **-27.8%**
165
+ - reasoning tokens: approximately **-27.8%**
166
+ - completion tokens: approximately **-29.2%**
167
+
168
+ The main benefit in this test was not higher raw per-token throughput.
169
+ Swift reached a similarly useful diagnostic answer with substantially fewer
170
+ reasoning and completion tokens.
171
+
172
+ This is a local workload test and should not be interpreted as a standardized
173
+ quality benchmark.
174
+
175
+ ## SGLang usage
176
+
177
+ ### Basic serving
178
+
179
+ ```bash
180
+ python -m sglang.launch_server \
181
+ --model-path /path/to/Swift-Qwen3.8-27B-FP8 \
182
+ --served-model-name Swift-Qwen3.8-27B-FP8 \
183
+ --host 0.0.0.0 \
184
+ --port 30000 \
185
+ --attention-backend flashinfer \
186
+ --reasoning-parser qwen3 \
187
+ --tool-call-parser qwen3_coder
188
+ ```
189
+
190
+ ### Native NEXTN / MTP speculative decoding
191
+
192
+ The retained native MTP head can be used with SGLang:
193
+
194
+ ```bash
195
+ python -m sglang.launch_server \
196
+ --model-path /path/to/Swift-Qwen3.8-27B-FP8 \
197
+ --served-model-name Swift-Qwen3.8-27B-FP8 \
198
+ --host 0.0.0.0 \
199
+ --port 30000 \
200
+ --attention-backend flashinfer \
201
+ --reasoning-parser qwen3 \
202
+ --tool-call-parser qwen3_coder \
203
+ --speculative-algorithm NEXTN \
204
+ --speculative-num-steps 3 \
205
+ --speculative-eagle-topk 1 \
206
+ --speculative-num-draft-tokens 4
207
+ ```
208
+
209
+ In the validated SGLang build, NEXTN is internally represented through the
210
+ EAGLE speculative-decoding path.
211
+
212
+ ### DFlash2 speculative decoding
213
+
214
+ Best local throughput was obtained with a compatible external DFlash2 draft
215
+ checkpoint:
216
+
217
+ ```bash
218
+ python -m sglang.launch_server \
219
+ --model-path /path/to/Swift-Qwen3.8-27B-FP8 \
220
+ --served-model-name Swift-Qwen3.8-27B-FP8 \
221
+ --host 0.0.0.0 \
222
+ --port 30000 \
223
+ --attention-backend flashinfer \
224
+ --reasoning-parser qwen3 \
225
+ --tool-call-parser qwen3_coder \
226
+ --speculative-algorithm DFLASH \
227
+ --speculative-draft-model-path /path/to/Qwen3.8-27B-DFlash2 \
228
+ --speculative-num-draft-tokens 8
229
+ ```
230
+
231
+ The DFlash2 weights are not redistributed in this repository.
232
+
233
+ Because speculative decoding verifies proposed tokens against the target
234
+ model, the external draft model affects acceptance rate and speed rather than
235
+ replacing the target model's token distribution.
236
+
237
+ ## Notes
238
+
239
+ This repository contains the quantized target checkpoint only.
240
+
241
+ It does not include:
242
+
243
+ - a DFlash2 draft checkpoint
244
+ - the original BF16 Swift checkpoint
245
+ - SGLang runtime binaries or containers
246
+
247
+ Performance depends heavily on GPU architecture, SGLang version, attention
248
+ backend, context length, concurrency, KV-cache configuration and speculative
249
+ decoding parameters.
250
 
251
  ## License
252
 
253
+ This checkpoint is derived from:
254
+
255
+ [`ukisai/Swift-Qwen3.8-27b`](https://huggingface.co/ukisai/Swift-Qwen3.8-27b)
256
+
257
+ and follows the **Swift Open License v1.0** applicable to the upstream model.
258
+
259
+ Please refer to the upstream repository and its license text for the
260
+ authoritative licensing terms.
261
+
262
+ No additional rights to the upstream model are granted by this repository.
263
+
264
+ ## Attribution
265
+
266
+ Original model:
267
+
268
+ - UkisAI
269
+ - `ukisai/Swift-Qwen3.8-27b`
270
+
271
+ FP8 conversion, validation and local performance measurements for this
272
+ repository were performed independently by the repository maintainer.
273
+
274
+ ## Citation
275
 
276
+ For the underlying Swift model, please cite or reference the upstream project:
277
 
278
+ ```bibtex
279
+ @misc{swift-qwen3.8-27b,
280
+ title = {Swift-Qwen3.8-27B},
281
+ author = {UkisAI},
282
+ year = {2026},
283
+ url = {https://huggingface.co/ukisai/Swift-Qwen3.8-27b}
284
+ }
285
+ ```