plunderstruck commited on
Commit
c75438f
Β·
verified Β·
1 Parent(s): 27ddbbb

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +11 -29
README.md CHANGED
@@ -108,13 +108,12 @@ Experimental **AMD Strix Halo (gfx1151)** quant of [**microsoft/FastContext-1.0-
108
  <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th>
109
  </tr></thead>
110
  <tbody>
111
- <tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-COHERENT-embF16.gguf</code> β˜…</td><td style="border:1px solid currentColor; padding:7px 10px;">all-dual</td><td style="border:1px solid currentColor; padding:7px 10px;">2.8 GB</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>recommended</b> β€” lowest measured KL vs BF16 (Β§04)</td></tr>
112
- <tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-embF16-imatrix.gguf</code></td><td style="border:1px solid currentColor; padding:7px 10px;">fast</td><td style="border:1px solid currentColor; padding:7px 10px;">2.7 GB</td><td style="border:1px solid currentColor; padding:7px 10px;">~same fidelity, slightly smaller/faster</td></tr>
113
  </tbody>
114
  </table>
115
  </div>
116
 
117
- Both share genuine **f16 embeddings** (from BF16) + the code-weighted imatrix (see Β§04). The **COHERENT** build (β˜…) puts every body tensor on the **dual-scale** `q4_0_rocmfp4` kernel β€” lowest measured KL vs the BF16 reference at ~the same decode speed β€” vs the STRIX build's faster single-scale `q4_0_rocmfp4_fast` bulk. The Qwen (ChatML) chat template is **baked into the GGUF** β€” just pass `--jinja`.
118
 
119
  <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
120
  <b>NOTE // TIED EMBEDDINGS.</b> FastContext has <code>tie_word_embeddings=True</code>, so there's <b>no separate output head</b> β€” the token-embedding tensor doubles as the lm-head. Setting <code>--token-embedding-type f16</code> therefore gives an <b>f16 embedding <i>and</i> f16 output head</b> in one (no <code>headQ6</code> variant needed β€” f16 already beats Q6 there).
@@ -127,7 +126,7 @@ Run from the folder holding the `.gguf` (the Qwen ChatML template is baked in
127
  ```bash
128
  env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
129
  llama-server \
130
- -m FastContext-1.0-4B-SFT-ROCmFP4-COHERENT-embF16.gguf \
131
  --alias fastcontext-4b \
132
  --host 0.0.0.0 \
133
  --port 8080 \
@@ -199,36 +198,23 @@ FastContext isn't a general chat model β€” it's a **repository-exploration subag
199
  <tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">DECODE Β· short context</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~68 t/s (Vulkan / Ryzen AI Max+ 395)</td></tr>
200
  <tr><td style="border:1px solid currentColor; padding:8px 11px;">SPECULATIVE DECODE</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">none (no MTP head)</td></tr>
201
  <tr><td style="border:1px solid currentColor; padding:8px 11px;">CONTEXT</td><td style="border:1px solid currentColor; padding:8px 11px;">256K native (dense attention)</td></tr>
202
- <tr><td style="border:1px solid currentColor; padding:8px 11px;">QUANTIZATION</td><td style="border:1px solid currentColor; padding:8px 11px;">COHERENT body + imatrix (measured win β€” below)</td></tr>
203
  </tbody>
204
  </table>
205
  </div>
206
 
207
- **Recommended build = COHERENT (we measured it).** Both builds use f16 tied emb/head + the same imatrix; the lever swept here is the **body kernel**, ranked by **KL divergence vs the true BF16** on held-out code (lower = more faithful). The **all-dual-scale body** (COHERENT) beats the fast-body STRIX build on **every** metric at ~the same decode speed:
208
 
209
- <div style="overflow:hidden; border-radius:0;">
210
- <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
211
- <thead><tr>
212
- <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Build (imatrix + embF16, tied head)</th>
213
- <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Body</th>
214
- <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Mean KLD vs BF16 ↓</th>
215
- <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Median KLD ↓</th>
216
- <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Top-token</th>
217
- <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">PPL(Q) ↓</th>
218
- </tr></thead>
219
- <tbody>
220
- <tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>COHERENT</code> β˜…</td><td style="border:1px solid currentColor; padding:7px 10px;">all-dual</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>0.03422</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>0.00955</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>92.08%</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>4.192</b></td></tr>
221
- <tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>STRIX</code></td><td style="border:1px solid currentColor; padding:7px 10px;">fast</td><td style="border:1px solid currentColor; padding:7px 10px;">0.03934</td><td style="border:1px solid currentColor; padding:7px 10px;">0.01016</td><td style="border:1px solid currentColor; padding:7px 10px;">91.38%</td><td style="border:1px solid currentColor; padding:7px 10px;">4.213</td></tr>
222
- </tbody>
223
- </table>
224
- </div>
225
 
226
- A **clean sweep**: COHERENT is lower on mean KLD (βˆ’13%), median KLD (βˆ’6%), RMS Ξ”p (6.43% vs 6.95%), **and** perplexity (4.192 vs 4.213), and higher on same-top-token (+0.70 pp) β€” every metric, same direction (BF16 reference PPL 4.074). So it's the default; STRIX stays as a marginally smaller/faster fallback.
 
 
227
 
228
  **Fast on its own.** ~68 t/s short-context decode on a Ryzen AI Max+ 395 (Vulkan0, measured `llama-bench tg128`). It's a 4B dense Qwen3 with **no MTP head**, so there's no speculative decoding β€” it doesn't need it, and at 4B it's a cheap explorer you can run several of in parallel.
229
 
230
  <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
231
- <b>NOTE // imatrix.</b> Both builds are quantized <b>with</b> an importance matrix (Kalomaze <code>groups_merged</code> + froggeric <code>code</code>/<code>technical</code>, via <a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a>), computed on this model's BF16. We measured the <b>COHERENT-vs-STRIX</b> comparison above (both imatrix); we did <b>not</b> run a separate imatrix-vs-no-imatrix ablation on this model. Scope: the KL/PPL figures are a fidelity-vs-BF16 measurement on a held-out code slice, <b>not</b> an absolute coding benchmark.
232
  </div>
233
 
234
  <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> Β· BUILD (REPRODUCIBLE)</div>
@@ -240,12 +226,8 @@ python convert_hf_to_gguf.py FastContext-1.0-4B-SFT/ --outtype bf16 --outfile Fa
240
  # 1) imatrix on the BF16 (general+code: Kalomaze groups_merged + froggeric code/technical)
241
  llama-imatrix -m FastContext-1.0-4B-SFT-BF16.gguf -f general+code-calib.txt -o fastcontext-4b.imatrix -c 512 -ngl 999
242
 
243
- # 2) RECOMMENDED: COHERENT all-dual body + f16 tied emb/head (the β˜… file) β€” lowest KL (Β§04).
244
  # tie_word_embeddings=True -> --token-embedding-type f16 also gives an f16 output head; no --output-tensor-type.
245
- llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
246
- FastContext-1.0-4B-SFT-BF16.gguf FastContext-1.0-4B-SFT-ROCmFP4-COHERENT-embF16.gguf Q4_0_ROCMFP4_COHERENT
247
-
248
- # fast-body STRIX fallback (same f16 emb + imatrix)
249
  llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
250
  FastContext-1.0-4B-SFT-BF16.gguf FastContext-1.0-4B-SFT-ROCmFP4-STRIX-embF16-imatrix.gguf Q4_0_ROCMFP4_STRIX
251
  ```
 
108
  <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th>
109
  </tr></thead>
110
  <tbody>
111
+ <tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-embF16-imatrix.gguf</code> β˜…</td><td style="border:1px solid currentColor; padding:7px 10px;">fast</td><td style="border:1px solid currentColor; padding:7px 10px;">2.7 GB</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>the one build</b> β€” best speed/quality balance: f16 tied embeddings/head on the fast single-scale body</td></tr>
 
112
  </tbody>
113
  </table>
114
  </div>
115
 
116
+ One file β€” the **best speed/quality balance** in ROCmFP4 for Strix Halo. It keeps the quality lever that's actually *felt* β€” genuine **f16 embeddings (from BF16), which also serve as the output head since the model ties them** β€” on the fast single-scale `q4_0_rocmfp4_fast` body + a code-weighted imatrix (see Β§04). The Qwen (ChatML) chat template is **baked into the GGUF** β€” just pass `--jinja`.
117
 
118
  <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
119
  <b>NOTE // TIED EMBEDDINGS.</b> FastContext has <code>tie_word_embeddings=True</code>, so there's <b>no separate output head</b> β€” the token-embedding tensor doubles as the lm-head. Setting <code>--token-embedding-type f16</code> therefore gives an <b>f16 embedding <i>and</i> f16 output head</b> in one (no <code>headQ6</code> variant needed β€” f16 already beats Q6 there).
 
126
  ```bash
127
  env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
128
  llama-server \
129
+ -m FastContext-1.0-4B-SFT-ROCmFP4-STRIX-embF16-imatrix.gguf \
130
  --alias fastcontext-4b \
131
  --host 0.0.0.0 \
132
  --port 8080 \
 
198
  <tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">DECODE Β· short context</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~68 t/s (Vulkan / Ryzen AI Max+ 395)</td></tr>
199
  <tr><td style="border:1px solid currentColor; padding:8px 11px;">SPECULATIVE DECODE</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">none (no MTP head)</td></tr>
200
  <tr><td style="border:1px solid currentColor; padding:8px 11px;">CONTEXT</td><td style="border:1px solid currentColor; padding:8px 11px;">256K native (dense attention)</td></tr>
201
+ <tr><td style="border:1px solid currentColor; padding:8px 11px;">QUANTIZATION</td><td style="border:1px solid currentColor; padding:8px 11px;">fast single-scale body + f16 tied emb/head + code-weighted imatrix</td></tr>
202
  </tbody>
203
  </table>
204
  </div>
205
 
206
+ **This is the best speed/quality balance in ROCmFP4 β€” by design, not the absolute fastest.** It keeps the one quality lever that's actually *felt* β€” genuine **f16 embeddings**, which on this model **double as the output head** (`tie_word_embeddings=True`), so a single f16 tensor sharpens both the input and output side at near-zero decode cost (it's a lookup, not a matmul) β€” on top of the fast single-scale `q4_0_rocmfp4_fast` body + a code-weighted imatrix. A leaner Q5-embedding build would shave a couple tok/s but degrades that lever; we keep full f16.
207
 
208
+ We didn't re-run the entire rocmfp4 lever sweep on this 4B. We ran it exhaustively on the larger **[Qwen3.6-27B](https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF)** β€” KL divergence vs the BF16 reference plus `llama-bench` decode across an all-dual-scale body, selective higher-precision tensors, and full f16 embeddings. The finding there: **an all-dual-scale body and selective higher-precision tensors both cost decode speed for a KL improvement that sat inside the measurement noise**, so the fast single-scale body + f16 embeddings is the balance point. That conclusion carries to FastContext β€” same format, same kernels β€” so we ship the one build that lands on it rather than a slower variant that wins KL only inside the noise.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
209
 
210
+ <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.9;">
211
+ <b>WANT MAXIMUM FIDELITY INSTEAD OF SPEED?</b> Grab a <b>Q6_K / Q8 GGUF of the base</b> from <a href="https://huggingface.co/microsoft/FastContext-1.0-4B-SFT"><b>microsoft/FastContext-1.0-4B-SFT</b></a> β€” higher-bit GGUFs run on this same fork. We optimize for throughput in ROCmFP4; if you want the last bit of fidelity over speed, a Q6_K/Q8 of the base is the one to grab.
212
+ </div>
213
 
214
  **Fast on its own.** ~68 t/s short-context decode on a Ryzen AI Max+ 395 (Vulkan0, measured `llama-bench tg128`). It's a 4B dense Qwen3 with **no MTP head**, so there's no speculative decoding β€” it doesn't need it, and at 4B it's a cheap explorer you can run several of in parallel.
215
 
216
  <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
217
+ <b>NOTE // imatrix.</b> This build is quantized <b>with</b> an importance matrix (Kalomaze <code>groups_merged</code> + froggeric <code>code</code>/<code>technical</code>, via <a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a>), computed on this model's BF16. We did <b>not</b> run a separate imatrix-vs-no-imatrix ablation on this 4B; at 4+ bpw imatrix is a free polish, not a transformation. Scope note: any fidelity-vs-BF16 figures are a held-out measurement, <b>not</b> an absolute coding benchmark.
218
  </div>
219
 
220
  <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> Β· BUILD (REPRODUCIBLE)</div>
 
226
  # 1) imatrix on the BF16 (general+code: Kalomaze groups_merged + froggeric code/technical)
227
  llama-imatrix -m FastContext-1.0-4B-SFT-BF16.gguf -f general+code-calib.txt -o fastcontext-4b.imatrix -c 512 -ngl 999
228
 
229
+ # 2) THE ONE BUILD: fast single-scale STRIX body + f16 tied emb/head + imatrix (the β˜… file) β€” the balance point (Β§04).
230
  # tie_word_embeddings=True -> --token-embedding-type f16 also gives an f16 output head; no --output-tensor-type.
 
 
 
 
231
  llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
232
  FastContext-1.0-4B-SFT-BF16.gguf FastContext-1.0-4B-SFT-ROCmFP4-STRIX-embF16-imatrix.gguf Q4_0_ROCMFP4_STRIX
233
  ```