File size: 23,800 Bytes
2fd4bb5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
---
base_model: microsoft/FastContext-1.0-4B-SFT
license: mit
library_name: gguf
tags:
- gguf
- rocmfp4
- qwen3
- fastcontext
- subagent
- repository-exploration
- coder
- agentic
- imatrix
- strix-halo
- amd
- rocm
- vulkan
language:
- en
base_model_relation: quantized
---

<div style="border:2px solid currentColor; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;">
  <div style="border-bottom:1px solid currentColor; padding:6px 12px; font-size:11px; letter-spacing:3px; text-transform:uppercase; opacity:0.7; text-align:center;">PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO · gfx1151</div>
  <div style="padding:14px; display:flex; flex-wrap:wrap; align-items:center; justify-content:center; gap:18px;">
<pre style="margin:0; flex:0 0 auto; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,monospace; font-size:5px; line-height:1.1; letter-spacing:0;">
            ▗▇▇▇▇▇▇▇▖                 
           ▗█▘▝██████▖                
          ▗▛   ▝██████▆▆▆▆▆▆▆▆▆▆▅     
         ▟▛    ▗█████████████████▙▖   
   ▄▄▄▄▄▟▛    ▟████████████████████▖  
 ▗██▌    ▚▖   ▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔█▘  
▗████▖    ▜▖                    ▗█▘   
▜█████▙    ▜▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▀▀▀▀▀▜▙    
 ▜█████▙    ▝████████████▛       ▜▙   
  ▜█████▙    ▝██████████▛    ▃    ▜▙  
   ▀█████▙▖   ▝████████▘    ▟█▙    ▀▙ 
    ▝██████▖   ▝▜█████▘    ▟███▙▂▂▂▂▐█
    ▟███████▖    ▜███▘   ▗███████████▛
   ▟█████████▄    ▜▛    ▗███████████▀ 
  ▝█████▀        ▗▛    ▗██████▀▀▀▀▀▘  
    ▜██▘        ▗▛    ▟█████▛▘        
     ▜█▇▇▇▇▇▇▇▇▇█▖   ▟█████▛          
                ▝█▖ ▟█████▛           
                 ▝███████▀            
</pre>
    <div style="flex:0 1 auto; max-width:100%; text-align:center;">
      <div style="font-size:23px; font-weight:800; letter-spacing:1px;">FASTCONTEXT-1.0-4B</div>
      <div style="font-size:12.5px; letter-spacing:1px; opacity:0.8; margin-top:5px;"><span style="white-space:nowrap;">4-BIT ROCmFP4</span> · <span style="white-space:nowrap;">QWEN3 DENSE 4B</span> · <span style="white-space:nowrap;">REPO-EXPLORATION SUBAGENT</span> · <span style="white-space:nowrap;">CODE-WEIGHTED IMATRIX</span> · <span style="white-space:nowrap;">SINGLE AMD APU</span></div>
    </div>
  </div>
  <table style="display:table; table-layout:fixed; width:100%; margin:0; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">
    <tr>
      <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">FORMAT</div><div style="font-weight:700;">ROCmFP4 4-BIT</div></td>
      <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PRECISION</div><div style="font-weight:700;">~4.5 BPW</div></td>
      <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">ARCH</div><div style="font-weight:700;">QWEN3 DENSE</div></td>
      <td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CONTEXT</div><div style="font-weight:700;">256 K</div></td>
    </tr>
    <tr>
      <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PARAMS</div><div style="font-weight:700;">4B DENSE</div></td>
      <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">DRAFT</div><div style="font-weight:700;">NO MTP</div></td>
      <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">BACKEND</div><div style="font-weight:700;">VULKAN0</div></td>
      <td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">LICENSE</div><div style="font-weight:700;">MIT</div></td>
    </tr>
  </table>
</div>

<div style="border:2px solid #dc2626; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;">
<b style="color:#dc2626; letter-spacing:1px;">⚠ REQUIRES THE ROCmFP4 FORK</b><br>
The custom <code>q4_0_rocmfp4</code> / <code>q4_0_rocmfp4_fast</code> tensor types <b>will not load in stock llama.cpp, LM Studio, or Ollama</b>. Build/run with <a href="https://github.com/charlie12345/rocmfp4-llama">charlie12345/rocmfp4-llama</a> · branch <code>mtp-rocmfp4-strix</code>.
</div>

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;">
<b>NOTE //</b> Ignore HuggingFace's auto-detected "F16"/16-bit badge — its parser can't read ROCmFP4 and mislabels the file. These are <b>~4.5 bpw 4-bit</b> ROCmFP4 files; pick by filename in <i>Files and versions</i>.
</div>

Experimental **AMD Strix Halo (gfx1151)** quant of [**microsoft/FastContext-1.0-4B-SFT**](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT) — Microsoft's **repository-exploration subagent** for coding agents. Instead of one model both exploring the repo and solving the task, FastContext is invoked on demand by a main agent, fires **parallel read-only tool calls** (READ / GLOB / GREP), and returns **compact file paths + line ranges** as focused context. Architecturally it's a plain **Qwen3 dense 4B** (`Qwen3ForCausalLM`, 36 layers, hidden 2560, 256K context, MIT-licensed), here in the custom **ROCmFP4** 4-bit format, **imatrix-quantized**.

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">01</span> · FILES</div>

<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<thead><tr>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">File</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Body</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Size</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th>
</tr></thead>
<tbody>
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-COHERENT-embF16.gguf</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">all-dual</td><td style="border:1px solid currentColor; padding:7px 10px;">2.8 GB</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>recommended</b> — lowest measured KL vs BF16 (§04)</td></tr>
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-embF16-imatrix.gguf</code></td><td style="border:1px solid currentColor; padding:7px 10px;">fast</td><td style="border:1px solid currentColor; padding:7px 10px;">2.7 GB</td><td style="border:1px solid currentColor; padding:7px 10px;">~same fidelity, slightly smaller/faster</td></tr>
</tbody>
</table>
</div>

Both share genuine **f16 embeddings** (from BF16) + the code-weighted imatrix (see §04). The **COHERENT** build (★) puts every body tensor on the **dual-scale** `q4_0_rocmfp4` kernel — lowest measured KL vs the BF16 reference at ~the same decode speed — vs the STRIX build's faster single-scale `q4_0_rocmfp4_fast` bulk. The Qwen (ChatML) chat template is **baked into the GGUF** — just pass `--jinja`.

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
<b>NOTE // TIED EMBEDDINGS.</b> FastContext has <code>tie_word_embeddings=True</code>, so there's <b>no separate output head</b> — the token-embedding tensor doubles as the lm-head. Setting <code>--token-embedding-type f16</code> therefore gives an <b>f16 embedding <i>and</i> f16 output head</b> in one (no <code>headQ6</code> variant needed — f16 already beats Q6 there).
</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">02</span> · QUICK START</div>

Run from the folder holding the `.gguf` (the Qwen ChatML template is baked in — just pass `--jinja`):

```bash
env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server \
  -m FastContext-1.0-4B-SFT-ROCmFP4-COHERENT-embF16.gguf \
  --alias fastcontext-4b \
  --host 0.0.0.0 \
  --port 8080 \
  -c 262144 \
  -ctk f16 \
  -ctv f16 \
  --temp 0.7 \
  --top-p 0.8 \
  --top-k 20 \
  -dev Vulkan0 \
  -ngl 999 \
  -fa on \
  -b 2048 \
  -ub 256 \
  -t 16 \
  -tb 16 \
  -cpent 256 \
  -ctxcp 32 \
  --cache-reuse 256 \
  --cache-ram 65536 \
  --jinja \
  --parallel 1 \
  --metrics \
  --no-mmap
```

<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">
<thead><tr>
<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px; width:40%;">Flag</th>
<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Function</th>
</tr></thead>
<tbody>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>HSA_OVERRIDE_GFX_VERSION=11.5.1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">treat the APU as gfx1151 (Strix Halo)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>GGML_HIP_ENABLE_UNIFIED_MEMORY=1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">allow use of the full 128 GB unified memory</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-dev Vulkan0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">run on Vulkan — fastest backend for ROCmFP4 on Strix Halo</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ngl 999 · -fa on</code></td><td style="border:1px solid currentColor; padding:6px 10px;">offload all layers · flash attention</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-c 262144</code></td><td style="border:1px solid currentColor; padding:6px 10px;">context length (256K)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-b 2048 · -ub 256 · -t/-tb 16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">prefill batch / micro-batch · CPU threads</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ctk f16 · -ctv f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">f16 KV cache — how we run it (cheap on a 4B); drop to <code>q8_0</code>/<code>q4_0</code> to use less memory at deep context</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-cpent · -ctxcp · --cache-reuse · --cache-ram 65536</code></td><td style="border:1px solid currentColor; padding:6px 10px;">cross-turn KV checkpointing + 64 GB resident reuse cache</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--temp 0.7 --top-p 0.8 --top-k 20</code></td><td style="border:1px solid currentColor; padding:6px 10px;">Qwen3 recommended sampling (instruct/non-thinking)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--jinja --parallel 1 --metrics --no-mmap</code></td><td style="border:1px solid currentColor; padding:6px 10px;">apply baked ChatML template · single slot · metrics · weights in RAM</td></tr>
</tbody>
</table>
</div>

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
<b>NOTE //</b> No <code>--spec-*</code> / <code>--spec-type draft-mtp</code> flags — this arch has <b>no MTP head</b> (see §04). It's already fast on its own.
</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">03</span> · USING IT AS A SUBAGENT</div>

FastContext isn't a general chat model — it's a **repository-exploration subagent** meant to be **called by your main coding agent**, not driven directly. The intended loop: the main agent delegates "find the relevant context for X" → FastContext issues **parallel read-only tool calls** (`READ`, `GLOB`, `GREP`) → returns **compact file paths + line ranges**, which the main agent folds into its own context to do the actual work. The point is to keep repo-exploration tokens *out* of the main agent's window.

- **Chat template:** Qwen (ChatML) is baked into the GGUF — just pass `--jinja`.
- **Tool calling:** it emits structured `READ`/`GLOB`/`GREP` calls — wire those tools into your harness and use a Qwen/Hermes-style tool-call parser so they're parsed rather than printed. **See the [upstream model card](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT) for the exact subagent protocol + tool schema** (it expects a specific invocation format).
- **Sampling:** temp `0.7`, top-p `0.8`, top-k `20` (Qwen3 instruct defaults) — already set in §02.

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
<b>NOTE //</b> It's small (4B) and fast (~68 t/s, §04) by design — a cheap, disposable explorer you can fan out in parallel next to a larger main model on the same box. The cross-turn reuse cache (<code>--cache-reuse</code> / <code>--cache-ram</code>) keeps repeated exploration over the same repo cheap.
</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">04</span> · PERFORMANCE &amp; QUALITY</div>

<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<tbody>
<tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">DECODE · short context</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~68 t/s (Vulkan / Ryzen AI Max+ 395)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">SPECULATIVE DECODE</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">none (no MTP head)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">CONTEXT</td><td style="border:1px solid currentColor; padding:8px 11px;">256K native (dense attention)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">QUANTIZATION</td><td style="border:1px solid currentColor; padding:8px 11px;">COHERENT body + imatrix (measured win — below)</td></tr>
</tbody>
</table>
</div>

**Recommended build = COHERENT (we measured it).** Both builds use f16 tied emb/head + the same imatrix; the lever swept here is the **body kernel**, ranked by **KL divergence vs the true BF16** on held-out code (lower = more faithful). The **all-dual-scale body** (COHERENT) beats the fast-body STRIX build on **every** metric at ~the same decode speed:

<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<thead><tr>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Build (imatrix + embF16, tied head)</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Body</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Mean KLD vs BF16 ↓</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Median KLD ↓</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Top-token</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">PPL(Q) ↓</th>
</tr></thead>
<tbody>
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>COHERENT</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">all-dual</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>0.03422</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>0.00955</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>92.08%</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>4.192</b></td></tr>
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>STRIX</code></td><td style="border:1px solid currentColor; padding:7px 10px;">fast</td><td style="border:1px solid currentColor; padding:7px 10px;">0.03934</td><td style="border:1px solid currentColor; padding:7px 10px;">0.01016</td><td style="border:1px solid currentColor; padding:7px 10px;">91.38%</td><td style="border:1px solid currentColor; padding:7px 10px;">4.213</td></tr>
</tbody>
</table>
</div>

A **clean sweep**: COHERENT is lower on mean KLD (−13%), median KLD (−6%), RMS Δp (6.43% vs 6.95%), **and** perplexity (4.192 vs 4.213), and higher on same-top-token (+0.70 pp) — every metric, same direction (BF16 reference PPL 4.074). So it's the default; STRIX stays as a marginally smaller/faster fallback.

**Fast on its own.** ~68 t/s short-context decode on a Ryzen AI Max+ 395 (Vulkan0, measured `llama-bench tg128`). It's a 4B dense Qwen3 with **no MTP head**, so there's no speculative decoding — it doesn't need it, and at 4B it's a cheap explorer you can run several of in parallel.

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
<b>NOTE // imatrix.</b> Both builds are quantized <b>with</b> an importance matrix (Kalomaze <code>groups_merged</code> + froggeric <code>code</code>/<code>technical</code>, via <a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a>), computed on this model's BF16. We measured the <b>COHERENT-vs-STRIX</b> comparison above (both imatrix); we did <b>not</b> run a separate imatrix-vs-no-imatrix ablation on this model. Scope: the KL/PPL figures are a fidelity-vs-BF16 measurement on a held-out code slice, <b>not</b> an absolute coding benchmark.
</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> · BUILD (REPRODUCIBLE)</div>

```bash
# 0) convert the safetensors -> BF16 GGUF (plain qwen3 dense; no MTP, tied embeddings)
python convert_hf_to_gguf.py FastContext-1.0-4B-SFT/ --outtype bf16 --outfile FastContext-1.0-4B-SFT-BF16.gguf

# 1) imatrix on the BF16 (general+code: Kalomaze groups_merged + froggeric code/technical)
llama-imatrix -m FastContext-1.0-4B-SFT-BF16.gguf -f general+code-calib.txt -o fastcontext-4b.imatrix -c 512 -ngl 999

# 2) RECOMMENDED: COHERENT all-dual body + f16 tied emb/head (the ★ file) — lowest KL (§04).
#    tie_word_embeddings=True -> --token-embedding-type f16 also gives an f16 output head; no --output-tensor-type.
llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
  FastContext-1.0-4B-SFT-BF16.gguf  FastContext-1.0-4B-SFT-ROCmFP4-COHERENT-embF16.gguf  Q4_0_ROCMFP4_COHERENT

# fast-body STRIX fallback (same f16 emb + imatrix)
llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
  FastContext-1.0-4B-SFT-BF16.gguf  FastContext-1.0-4B-SFT-ROCmFP4-STRIX-embF16-imatrix.gguf  Q4_0_ROCMFP4_STRIX
```

> Experimental research build for AMD Strix Halo — hardware/driver/prompt-sensitive, may not reproduce elsewhere. Not native FP4 tensor-core execution.

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">06</span> · LINEAGE &amp; CREDITS</div>

<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<tbody>
<tr><td style="border:1px solid currentColor; padding:8px 11px; width:26%;">BASE MODEL</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/microsoft/FastContext-1.0-4B-SFT">microsoft/FastContext-1.0-4B-SFT</a> (MIT, Microsoft) · repository-exploration subagent · Qwen3 dense 4B (<code>Qwen3ForCausalLM</code>)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">CALIBRATION</td><td style="border:1px solid currentColor; padding:8px 11px;">Kalomaze <code>groups_merged</code> + froggeric <code>code</code>/<code>technical</code> via <a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a></td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">FORMAT + RUNTIME</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://github.com/charlie12345/rocmfp4-llama">charlie12345/rocmfp4-llama</a> (based on llama.cpp, MIT)</td></tr>
</tbody>
</table>
</div>

*Derivative quantization — verify the base model's license before redistribution / use.*