bjonor commited on
Commit
cbbb052
·
verified ·
1 Parent(s): 21055aa

Model card: surface SergiioB cookbook link (code row + patch notes)

Browse files
Files changed (1) hide show
  1. README.md +4 -2
README.md CHANGED
@@ -55,7 +55,7 @@ models below continue to apply (see [License](#license-and-attribution)).
55
  | Source revision | `048328f4059015b63f860a453bf94834af0db683` |
56
  | Calibration | `HuggingFaceH4/ultrachat_200k` `train_sft[:256]`, truncated to 2048 tokens |
57
  | Calibration revision | `8049631c405ae6576f93f445c6b8166f76f5505a` |
58
- | Code | quantization, verification and Intel-XPU serving recipe: [BjornNordblom/intel-arc-b70-quant](https://github.com/BjornNordblom/intel-arc-b70-quant) |
59
 
60
  The `quantize_config.json` is field-for-field identical to the community
61
  reference artifact
@@ -113,7 +113,9 @@ Notes for this base model on XPU:
113
 
114
  - **MTP draft must be built unquantized.** The checkpoint flags this via the
115
  `dynamic` exclusion, but the XPU build tested here also needs the draft layer
116
- built without `quant_config` (upstream patch used by [the recipe
 
 
117
  repo](https://github.com/BjornNordblom/intel-arc-b70-quant):
118
  `B70_MTP_BF16_DRAFT=1` gate plus a small metadata patch for the
119
  max-model-length boundary).
 
55
  | Source revision | `048328f4059015b63f860a453bf94834af0db683` |
56
  | Calibration | `HuggingFaceH4/ultrachat_200k` `train_sft[:256]`, truncated to 2048 tokens |
57
  | Calibration revision | `8049631c405ae6576f93f445c6b8166f76f5505a` |
58
+ | Code | quantization, verification and Intel-XPU serving recipe: [BjornNordblom/intel-arc-b70-quant](https://github.com/BjornNordblom/intel-arc-b70-quant); serving patches from [SergiioB/intel-arc-pro-b70-inference-cookbook](https://github.com/SergiioB/intel-arc-pro-b70-inference-cookbook) |
59
 
60
  The `quantize_config.json` is field-for-field identical to the community
61
  reference artifact
 
113
 
114
  - **MTP draft must be built unquantized.** The checkpoint flags this via the
115
  `dynamic` exclusion, but the XPU build tested here also needs the draft layer
116
+ built without `quant_config` (patch from [SergiioB's Intel Arc Pro B70
117
+ cookbook](https://github.com/SergiioB/intel-arc-pro-b70-inference-cookbook),
118
+ vendored in [the recipe
119
  repo](https://github.com/BjornNordblom/intel-arc-b70-quant):
120
  `B70_MTP_BF16_DRAFT=1` gate plus a small metadata patch for the
121
  max-model-length boundary).