bjonor commited on
Commit
21055aa
·
verified ·
1 Parent(s): 2fd81f4

Model card: link code repo, SergiioB cookbook (correct patch attribution) and reference artifact

Browse files
Files changed (1) hide show
  1. README.md +16 -7
README.md CHANGED
@@ -55,11 +55,12 @@ models below continue to apply (see [License](#license-and-attribution)).
55
  | Source revision | `048328f4059015b63f860a453bf94834af0db683` |
56
  | Calibration | `HuggingFaceH4/ultrachat_200k` `train_sft[:256]`, truncated to 2048 tokens |
57
  | Calibration revision | `8049631c405ae6576f93f445c6b8166f76f5505a` |
58
- | Repository | quantization recipe + verification: https://github.com/BjornNordblom/intel-arc-b70-quant |
59
 
60
  The `quantize_config.json` is field-for-field identical to the community
61
- reference artifact `SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16` except for
62
- the quant-time-only `meta.offload_to_disk` flag. Note that the reference artifact
 
63
  quantizes the **base** Qwen3.8-27B; this repository quantizes the **Swift**
64
  fine-tune.
65
 
@@ -112,7 +113,8 @@ Notes for this base model on XPU:
112
 
113
  - **MTP draft must be built unquantized.** The checkpoint flags this via the
114
  `dynamic` exclusion, but the XPU build tested here also needs the draft layer
115
- built without `quant_config` (upstream patch used by the recipe repo:
 
116
  `B70_MTP_BF16_DRAFT=1` gate plus a small metadata patch for the
117
  max-model-length boundary).
118
  - **`vllm-xpu-kernels` < 0.1.14.1 has a mixed-batch limitation**: batching
@@ -124,6 +126,11 @@ Notes for this base model on XPU:
124
  backport used by the recipe repo (a single-function change in `vllm/_xpu_ops.py`).
125
  - `--kv-cache-dtype fp8` is a serving choice, not part of the checkpoint.
126
 
 
 
 
 
 
127
  ### Transformers
128
 
129
  GPTQ checkpoints need a GPTQ-capable loader (`gptqmodel` or `auto-gptq`) and
@@ -139,7 +146,7 @@ MTP 3 speculative tokens, fp8 KV cache):
139
  |---|---|
140
  | Quantization contract (`verify_quant.py`) | PASS — bits 4, group 128, sym, `desc_act=false`, 15 MTP tensors preserved, 333 vision tensors, 400 modules quantized, `lm_head` untouched |
141
  | Tiny-model smoke test + endpoint/streaming test | PASS |
142
- | Decode, p512/g128, median of 5 | 58.8 tok/s (this artifact) vs 60.4 tok/s (`SergiioB/…` reference quant) |
143
  | MTP draft acceptance | 3.38 / 4 tokens accepted (59.5%) vs 3.43 (60.9%) for the reference quant |
144
  | Solo TTFT / decode (short prompt) | ~0.95 s / ~59 tok/s |
145
  | Concurrency | 12-request storm with 4 × ~10k-token prefills: 12/12 HTTP 200, no engine failure, MTP acceptance ~62% (with the mixed-batch fix above) |
@@ -247,8 +254,10 @@ Addendum for this quantization (the uploader's modification notice):
247
  - UkisAI for the Swift fine-tune and the licence terms above.
248
  - Alibaba Cloud / Qwen for Qwen3.8-27B (Apache-2.0).
249
  - The `gptqmodel` project for the quantizer.
250
- - Intel's Arc Pro B70 XPU inference cookbook for the serving patches (MTP BF16
251
- draft, GDN boundary handling, mixed-batch split-dispatch backport).
 
 
252
 
253
  ## Model card contact
254
 
 
55
  | Source revision | `048328f4059015b63f860a453bf94834af0db683` |
56
  | Calibration | `HuggingFaceH4/ultrachat_200k` `train_sft[:256]`, truncated to 2048 tokens |
57
  | Calibration revision | `8049631c405ae6576f93f445c6b8166f76f5505a` |
58
+ | Code | quantization, verification and Intel-XPU serving recipe: [BjornNordblom/intel-arc-b70-quant](https://github.com/BjornNordblom/intel-arc-b70-quant) |
59
 
60
  The `quantize_config.json` is field-for-field identical to the community
61
+ reference artifact
62
+ [`SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16`](https://huggingface.co/SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16)
63
+ except for the quant-time-only `meta.offload_to_disk` flag. Note that the reference artifact
64
  quantizes the **base** Qwen3.8-27B; this repository quantizes the **Swift**
65
  fine-tune.
66
 
 
113
 
114
  - **MTP draft must be built unquantized.** The checkpoint flags this via the
115
  `dynamic` exclusion, but the XPU build tested here also needs the draft layer
116
+ built without `quant_config` (upstream patch used by [the recipe
117
+ repo](https://github.com/BjornNordblom/intel-arc-b70-quant):
118
  `B70_MTP_BF16_DRAFT=1` gate plus a small metadata patch for the
119
  max-model-length boundary).
120
  - **`vllm-xpu-kernels` < 0.1.14.1 has a mixed-batch limitation**: batching
 
126
  backport used by the recipe repo (a single-function change in `vllm/_xpu_ops.py`).
127
  - `--kv-cache-dtype fp8` is a serving choice, not part of the checkpoint.
128
 
129
+ Full reproduction path — pinned environments, `quant_swift.py` /
130
+ `verify_quant.py`, `launch.sh` (MTP and vision flags), the XPU patches and the
131
+ benchmark harness — is in the
132
+ [recipe repository](https://github.com/BjornNordblom/intel-arc-b70-quant).
133
+
134
  ### Transformers
135
 
136
  GPTQ checkpoints need a GPTQ-capable loader (`gptqmodel` or `auto-gptq`) and
 
146
  |---|---|
147
  | Quantization contract (`verify_quant.py`) | PASS — bits 4, group 128, sym, `desc_act=false`, 15 MTP tensors preserved, 333 vision tensors, 400 modules quantized, `lm_head` untouched |
148
  | Tiny-model smoke test + endpoint/streaming test | PASS |
149
+ | Decode, p512/g128, median of 5 | 58.8 tok/s (this artifact) vs 60.4 tok/s ([`SergiioB/…`](https://huggingface.co/SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16) reference quant) |
150
  | MTP draft acceptance | 3.38 / 4 tokens accepted (59.5%) vs 3.43 (60.9%) for the reference quant |
151
  | Solo TTFT / decode (short prompt) | ~0.95 s / ~59 tok/s |
152
  | Concurrency | 12-request storm with 4 × ~10k-token prefills: 12/12 HTTP 200, no engine failure, MTP acceptance ~62% (with the mixed-batch fix above) |
 
254
  - UkisAI for the Swift fine-tune and the licence terms above.
255
  - Alibaba Cloud / Qwen for Qwen3.8-27B (Apache-2.0).
256
  - The `gptqmodel` project for the quantizer.
257
+ - SergiioB's [intel-arc-pro-b70-inference-cookbook](https://github.com/SergiioB/intel-arc-pro-b70-inference-cookbook)
258
+ for the Intel Arc Pro B70 XPU serving patches (MTP BF16 draft, GDN boundary
259
+ handling, mixed-batch split-dispatch backport); the copies vendored in the
260
+ recipe repo are MIT, Copyright (c) 2026 SergiioB.
261
 
262
  ## Model card contact
263