plunderstruck commited on
Commit
2fd4bb5
·
verified ·
1 Parent(s): 56c9b9c

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +247 -0
README.md ADDED
@@ -0,0 +1,247 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: microsoft/FastContext-1.0-4B-SFT
3
+ license: mit
4
+ library_name: gguf
5
+ tags:
6
+ - gguf
7
+ - rocmfp4
8
+ - qwen3
9
+ - fastcontext
10
+ - subagent
11
+ - repository-exploration
12
+ - coder
13
+ - agentic
14
+ - imatrix
15
+ - strix-halo
16
+ - amd
17
+ - rocm
18
+ - vulkan
19
+ language:
20
+ - en
21
+ base_model_relation: quantized
22
+ ---
23
+
24
+ <div style="border:2px solid currentColor; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;">
25
+ <div style="border-bottom:1px solid currentColor; padding:6px 12px; font-size:11px; letter-spacing:3px; text-transform:uppercase; opacity:0.7; text-align:center;">PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO · gfx1151</div>
26
+ <div style="padding:14px; display:flex; flex-wrap:wrap; align-items:center; justify-content:center; gap:18px;">
27
+ <pre style="margin:0; flex:0 0 auto; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,monospace; font-size:5px; line-height:1.1; letter-spacing:0;">
28
+ ▗▇▇▇▇▇▇▇▖
29
+ ▗█▘▝██████▖
30
+ ▗▛ ▝██████▆▆▆▆▆▆▆▆▆▆▅
31
+ ▟▛ ▗█████████████████▙▖
32
+ ▄▄▄▄▄▟▛ ▟████████████████████▖
33
+ ▗██▌ ▚▖ ▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔█▘
34
+ ▗████▖ ▜▖ ▗█▘
35
+ ▜█████▙ ▜▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▀▀▀▀▀▜▙
36
+ ▜█████▙ ▝████████████▛ ▜▙
37
+ ▜█████▙ ▝██████████▛ ▃ ▜▙
38
+ ▀█████▙▖ ▝████████▘ ▟█▙ ▀▙
39
+ ▝██████▖ ▝▜█████▘ ▟███▙▂▂▂▂▐█
40
+ ▟███████▖ ▜███▘ ▗███████████▛
41
+ ▟█████████▄ ▜▛ ▗███████████▀
42
+ ▝█████▀ ▗▛ ▗██████▀▀▀▀▀▘
43
+ ▜██▘ ▗▛ ▟█████▛▘
44
+ ▜█▇▇▇▇▇▇▇▇▇█▖ ▟█████▛
45
+ ▝█▖ ▟█████▛
46
+ ▝███████▀
47
+ </pre>
48
+ <div style="flex:0 1 auto; max-width:100%; text-align:center;">
49
+ <div style="font-size:23px; font-weight:800; letter-spacing:1px;">FASTCONTEXT-1.0-4B</div>
50
+ <div style="font-size:12.5px; letter-spacing:1px; opacity:0.8; margin-top:5px;"><span style="white-space:nowrap;">4-BIT ROCmFP4</span> · <span style="white-space:nowrap;">QWEN3 DENSE 4B</span> · <span style="white-space:nowrap;">REPO-EXPLORATION SUBAGENT</span> · <span style="white-space:nowrap;">CODE-WEIGHTED IMATRIX</span> · <span style="white-space:nowrap;">SINGLE AMD APU</span></div>
51
+ </div>
52
+ </div>
53
+ <table style="display:table; table-layout:fixed; width:100%; margin:0; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">
54
+ <tr>
55
+ <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">FORMAT</div><div style="font-weight:700;">ROCmFP4 4-BIT</div></td>
56
+ <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PRECISION</div><div style="font-weight:700;">~4.5 BPW</div></td>
57
+ <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">ARCH</div><div style="font-weight:700;">QWEN3 DENSE</div></td>
58
+ <td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CONTEXT</div><div style="font-weight:700;">256 K</div></td>
59
+ </tr>
60
+ <tr>
61
+ <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PARAMS</div><div style="font-weight:700;">4B DENSE</div></td>
62
+ <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">DRAFT</div><div style="font-weight:700;">NO MTP</div></td>
63
+ <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">BACKEND</div><div style="font-weight:700;">VULKAN0</div></td>
64
+ <td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">LICENSE</div><div style="font-weight:700;">MIT</div></td>
65
+ </tr>
66
+ </table>
67
+ </div>
68
+
69
+ <div style="border:2px solid #dc2626; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;">
70
+ <b style="color:#dc2626; letter-spacing:1px;">⚠ REQUIRES THE ROCmFP4 FORK</b><br>
71
+ The custom <code>q4_0_rocmfp4</code> / <code>q4_0_rocmfp4_fast</code> tensor types <b>will not load in stock llama.cpp, LM Studio, or Ollama</b>. Build/run with <a href="https://github.com/charlie12345/rocmfp4-llama">charlie12345/rocmfp4-llama</a> · branch <code>mtp-rocmfp4-strix</code>.
72
+ </div>
73
+
74
+ <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;">
75
+ <b>NOTE //</b> Ignore HuggingFace's auto-detected "F16"/16-bit badge — its parser can't read ROCmFP4 and mislabels the file. These are <b>~4.5 bpw 4-bit</b> ROCmFP4 files; pick by filename in <i>Files and versions</i>.
76
+ </div>
77
+
78
+ Experimental **AMD Strix Halo (gfx1151)** quant of [**microsoft/FastContext-1.0-4B-SFT**](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT) — Microsoft's **repository-exploration subagent** for coding agents. Instead of one model both exploring the repo and solving the task, FastContext is invoked on demand by a main agent, fires **parallel read-only tool calls** (READ / GLOB / GREP), and returns **compact file paths + line ranges** as focused context. Architecturally it's a plain **Qwen3 dense 4B** (`Qwen3ForCausalLM`, 36 layers, hidden 2560, 256K context, MIT-licensed), here in the custom **ROCmFP4** 4-bit format, **imatrix-quantized**.
79
+
80
+ <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">01</span> · FILES</div>
81
+
82
+ <div style="overflow:hidden; border-radius:0;">
83
+ <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
84
+ <thead><tr>
85
+ <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">File</th>
86
+ <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Body</th>
87
+ <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Size</th>
88
+ <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th>
89
+ </tr></thead>
90
+ <tbody>
91
+ <tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-COHERENT-embF16.gguf</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">all-dual</td><td style="border:1px solid currentColor; padding:7px 10px;">2.8 GB</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>recommended</b> — lowest measured KL vs BF16 (§04)</td></tr>
92
+ <tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-embF16-imatrix.gguf</code></td><td style="border:1px solid currentColor; padding:7px 10px;">fast</td><td style="border:1px solid currentColor; padding:7px 10px;">2.7 GB</td><td style="border:1px solid currentColor; padding:7px 10px;">~same fidelity, slightly smaller/faster</td></tr>
93
+ </tbody>
94
+ </table>
95
+ </div>
96
+
97
+ Both share genuine **f16 embeddings** (from BF16) + the code-weighted imatrix (see §04). The **COHERENT** build (★) puts every body tensor on the **dual-scale** `q4_0_rocmfp4` kernel — lowest measured KL vs the BF16 reference at ~the same decode speed — vs the STRIX build's faster single-scale `q4_0_rocmfp4_fast` bulk. The Qwen (ChatML) chat template is **baked into the GGUF** — just pass `--jinja`.
98
+
99
+ <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
100
+ <b>NOTE // TIED EMBEDDINGS.</b> FastContext has <code>tie_word_embeddings=True</code>, so there's <b>no separate output head</b> — the token-embedding tensor doubles as the lm-head. Setting <code>--token-embedding-type f16</code> therefore gives an <b>f16 embedding <i>and</i> f16 output head</b> in one (no <code>headQ6</code> variant needed — f16 already beats Q6 there).
101
+ </div>
102
+
103
+ <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">02</span> · QUICK START</div>
104
+
105
+ Run from the folder holding the `.gguf` (the Qwen ChatML template is baked in — just pass `--jinja`):
106
+
107
+ ```bash
108
+ env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
109
+ llama-server \
110
+ -m FastContext-1.0-4B-SFT-ROCmFP4-COHERENT-embF16.gguf \
111
+ --alias fastcontext-4b \
112
+ --host 0.0.0.0 \
113
+ --port 8080 \
114
+ -c 262144 \
115
+ -ctk f16 \
116
+ -ctv f16 \
117
+ --temp 0.7 \
118
+ --top-p 0.8 \
119
+ --top-k 20 \
120
+ -dev Vulkan0 \
121
+ -ngl 999 \
122
+ -fa on \
123
+ -b 2048 \
124
+ -ub 256 \
125
+ -t 16 \
126
+ -tb 16 \
127
+ -cpent 256 \
128
+ -ctxcp 32 \
129
+ --cache-reuse 256 \
130
+ --cache-ram 65536 \
131
+ --jinja \
132
+ --parallel 1 \
133
+ --metrics \
134
+ --no-mmap
135
+ ```
136
+
137
+ <div style="overflow:hidden; border-radius:0;">
138
+ <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">
139
+ <thead><tr>
140
+ <th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px; width:40%;">Flag</th>
141
+ <th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Function</th>
142
+ </tr></thead>
143
+ <tbody>
144
+ <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>HSA_OVERRIDE_GFX_VERSION=11.5.1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">treat the APU as gfx1151 (Strix Halo)</td></tr>
145
+ <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>GGML_HIP_ENABLE_UNIFIED_MEMORY=1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">allow use of the full 128 GB unified memory</td></tr>
146
+ <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-dev Vulkan0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">run on Vulkan — fastest backend for ROCmFP4 on Strix Halo</td></tr>
147
+ <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ngl 999 · -fa on</code></td><td style="border:1px solid currentColor; padding:6px 10px;">offload all layers · flash attention</td></tr>
148
+ <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-c 262144</code></td><td style="border:1px solid currentColor; padding:6px 10px;">context length (256K)</td></tr>
149
+ <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-b 2048 · -ub 256 · -t/-tb 16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">prefill batch / micro-batch · CPU threads</td></tr>
150
+ <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ctk f16 · -ctv f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">f16 KV cache — how we run it (cheap on a 4B); drop to <code>q8_0</code>/<code>q4_0</code> to use less memory at deep context</td></tr>
151
+ <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-cpent · -ctxcp · --cache-reuse · --cache-ram 65536</code></td><td style="border:1px solid currentColor; padding:6px 10px;">cross-turn KV checkpointing + 64 GB resident reuse cache</td></tr>
152
+ <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--temp 0.7 --top-p 0.8 --top-k 20</code></td><td style="border:1px solid currentColor; padding:6px 10px;">Qwen3 recommended sampling (instruct/non-thinking)</td></tr>
153
+ <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--jinja --parallel 1 --metrics --no-mmap</code></td><td style="border:1px solid currentColor; padding:6px 10px;">apply baked ChatML template · single slot · metrics · weights in RAM</td></tr>
154
+ </tbody>
155
+ </table>
156
+ </div>
157
+
158
+ <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
159
+ <b>NOTE //</b> No <code>--spec-*</code> / <code>--spec-type draft-mtp</code> flags — this arch has <b>no MTP head</b> (see §04). It's already fast on its own.
160
+ </div>
161
+
162
+ <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">03</span> · USING IT AS A SUBAGENT</div>
163
+
164
+ FastContext isn't a general chat model — it's a **repository-exploration subagent** meant to be **called by your main coding agent**, not driven directly. The intended loop: the main agent delegates "find the relevant context for X" → FastContext issues **parallel read-only tool calls** (`READ`, `GLOB`, `GREP`) → returns **compact file paths + line ranges**, which the main agent folds into its own context to do the actual work. The point is to keep repo-exploration tokens *out* of the main agent's window.
165
+
166
+ - **Chat template:** Qwen (ChatML) is baked into the GGUF — just pass `--jinja`.
167
+ - **Tool calling:** it emits structured `READ`/`GLOB`/`GREP` calls — wire those tools into your harness and use a Qwen/Hermes-style tool-call parser so they're parsed rather than printed. **See the [upstream model card](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT) for the exact subagent protocol + tool schema** (it expects a specific invocation format).
168
+ - **Sampling:** temp `0.7`, top-p `0.8`, top-k `20` (Qwen3 instruct defaults) — already set in §02.
169
+
170
+ <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
171
+ <b>NOTE //</b> It's small (4B) and fast (~68 t/s, §04) by design — a cheap, disposable explorer you can fan out in parallel next to a larger main model on the same box. The cross-turn reuse cache (<code>--cache-reuse</code> / <code>--cache-ram</code>) keeps repeated exploration over the same repo cheap.
172
+ </div>
173
+
174
+ <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">04</span> · PERFORMANCE &amp; QUALITY</div>
175
+
176
+ <div style="overflow:hidden; border-radius:0;">
177
+ <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
178
+ <tbody>
179
+ <tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">DECODE · short context</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~68 t/s (Vulkan / Ryzen AI Max+ 395)</td></tr>
180
+ <tr><td style="border:1px solid currentColor; padding:8px 11px;">SPECULATIVE DECODE</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">none (no MTP head)</td></tr>
181
+ <tr><td style="border:1px solid currentColor; padding:8px 11px;">CONTEXT</td><td style="border:1px solid currentColor; padding:8px 11px;">256K native (dense attention)</td></tr>
182
+ <tr><td style="border:1px solid currentColor; padding:8px 11px;">QUANTIZATION</td><td style="border:1px solid currentColor; padding:8px 11px;">COHERENT body + imatrix (measured win — below)</td></tr>
183
+ </tbody>
184
+ </table>
185
+ </div>
186
+
187
+ **Recommended build = COHERENT (we measured it).** Both builds use f16 tied emb/head + the same imatrix; the lever swept here is the **body kernel**, ranked by **KL divergence vs the true BF16** on held-out code (lower = more faithful). The **all-dual-scale body** (COHERENT) beats the fast-body STRIX build on **every** metric at ~the same decode speed:
188
+
189
+ <div style="overflow:hidden; border-radius:0;">
190
+ <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
191
+ <thead><tr>
192
+ <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Build (imatrix + embF16, tied head)</th>
193
+ <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Body</th>
194
+ <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Mean KLD vs BF16 ↓</th>
195
+ <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Median KLD ↓</th>
196
+ <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Top-token</th>
197
+ <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">PPL(Q) ↓</th>
198
+ </tr></thead>
199
+ <tbody>
200
+ <tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>COHERENT</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">all-dual</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>0.03422</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>0.00955</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>92.08%</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>4.192</b></td></tr>
201
+ <tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>STRIX</code></td><td style="border:1px solid currentColor; padding:7px 10px;">fast</td><td style="border:1px solid currentColor; padding:7px 10px;">0.03934</td><td style="border:1px solid currentColor; padding:7px 10px;">0.01016</td><td style="border:1px solid currentColor; padding:7px 10px;">91.38%</td><td style="border:1px solid currentColor; padding:7px 10px;">4.213</td></tr>
202
+ </tbody>
203
+ </table>
204
+ </div>
205
+
206
+ A **clean sweep**: COHERENT is lower on mean KLD (−13%), median KLD (−6%), RMS Δp (6.43% vs 6.95%), **and** perplexity (4.192 vs 4.213), and higher on same-top-token (+0.70 pp) — every metric, same direction (BF16 reference PPL 4.074). So it's the default; STRIX stays as a marginally smaller/faster fallback.
207
+
208
+ **Fast on its own.** ~68 t/s short-context decode on a Ryzen AI Max+ 395 (Vulkan0, measured `llama-bench tg128`). It's a 4B dense Qwen3 with **no MTP head**, so there's no speculative decoding — it doesn't need it, and at 4B it's a cheap explorer you can run several of in parallel.
209
+
210
+ <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
211
+ <b>NOTE // imatrix.</b> Both builds are quantized <b>with</b> an importance matrix (Kalomaze <code>groups_merged</code> + froggeric <code>code</code>/<code>technical</code>, via <a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a>), computed on this model's BF16. We measured the <b>COHERENT-vs-STRIX</b> comparison above (both imatrix); we did <b>not</b> run a separate imatrix-vs-no-imatrix ablation on this model. Scope: the KL/PPL figures are a fidelity-vs-BF16 measurement on a held-out code slice, <b>not</b> an absolute coding benchmark.
212
+ </div>
213
+
214
+ <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> · BUILD (REPRODUCIBLE)</div>
215
+
216
+ ```bash
217
+ # 0) convert the safetensors -> BF16 GGUF (plain qwen3 dense; no MTP, tied embeddings)
218
+ python convert_hf_to_gguf.py FastContext-1.0-4B-SFT/ --outtype bf16 --outfile FastContext-1.0-4B-SFT-BF16.gguf
219
+
220
+ # 1) imatrix on the BF16 (general+code: Kalomaze groups_merged + froggeric code/technical)
221
+ llama-imatrix -m FastContext-1.0-4B-SFT-BF16.gguf -f general+code-calib.txt -o fastcontext-4b.imatrix -c 512 -ngl 999
222
+
223
+ # 2) RECOMMENDED: COHERENT all-dual body + f16 tied emb/head (the ★ file) — lowest KL (§04).
224
+ # tie_word_embeddings=True -> --token-embedding-type f16 also gives an f16 output head; no --output-tensor-type.
225
+ llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
226
+ FastContext-1.0-4B-SFT-BF16.gguf FastContext-1.0-4B-SFT-ROCmFP4-COHERENT-embF16.gguf Q4_0_ROCMFP4_COHERENT
227
+
228
+ # fast-body STRIX fallback (same f16 emb + imatrix)
229
+ llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
230
+ FastContext-1.0-4B-SFT-BF16.gguf FastContext-1.0-4B-SFT-ROCmFP4-STRIX-embF16-imatrix.gguf Q4_0_ROCMFP4_STRIX
231
+ ```
232
+
233
+ > Experimental research build for AMD Strix Halo — hardware/driver/prompt-sensitive, may not reproduce elsewhere. Not native FP4 tensor-core execution.
234
+
235
+ <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">06</span> · LINEAGE &amp; CREDITS</div>
236
+
237
+ <div style="overflow:hidden; border-radius:0;">
238
+ <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
239
+ <tbody>
240
+ <tr><td style="border:1px solid currentColor; padding:8px 11px; width:26%;">BASE MODEL</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/microsoft/FastContext-1.0-4B-SFT">microsoft/FastContext-1.0-4B-SFT</a> (MIT, Microsoft) · repository-exploration subagent · Qwen3 dense 4B (<code>Qwen3ForCausalLM</code>)</td></tr>
241
+ <tr><td style="border:1px solid currentColor; padding:8px 11px;">CALIBRATION</td><td style="border:1px solid currentColor; padding:8px 11px;">Kalomaze <code>groups_merged</code> + froggeric <code>code</code>/<code>technical</code> via <a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a></td></tr>
242
+ <tr><td style="border:1px solid currentColor; padding:8px 11px;">FORMAT + RUNTIME</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://github.com/charlie12345/rocmfp4-llama">charlie12345/rocmfp4-llama</a> (based on llama.cpp, MIT)</td></tr>
243
+ </tbody>
244
+ </table>
245
+ </div>
246
+
247
+ *Derivative quantization — verify the base model's license before redistribution / use.*