GGUF
English
rocmfp4
qwen3
fastcontext
subagent
repository-exploration
coder
agentic
imatrix
strix-halo
amd
rocm
vulkan
conversational
Instructions to use plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
Use Docker
docker model run hf.co/plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
- LM Studio
- Jan
- Ollama
How to use plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF with Ollama:
ollama run hf.co/plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
- Unsloth Desktop
- Pi
How to use plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF with Docker Model Runner:
docker model run hf.co/plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
- Lemonade
How to use plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
Run and chat with the model
lemonade run user.FastContext-1.0-4B-SFT-ROCmFP4-GGUF-BF16
List all available models
lemonade list
- Hermes Agent
How to use plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "plunderstruck/FastContext-1.0-4B-SFT-ROCmFP4-GGUF:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,247 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: microsoft/FastContext-1.0-4B-SFT
|
| 3 |
+
license: mit
|
| 4 |
+
library_name: gguf
|
| 5 |
+
tags:
|
| 6 |
+
- gguf
|
| 7 |
+
- rocmfp4
|
| 8 |
+
- qwen3
|
| 9 |
+
- fastcontext
|
| 10 |
+
- subagent
|
| 11 |
+
- repository-exploration
|
| 12 |
+
- coder
|
| 13 |
+
- agentic
|
| 14 |
+
- imatrix
|
| 15 |
+
- strix-halo
|
| 16 |
+
- amd
|
| 17 |
+
- rocm
|
| 18 |
+
- vulkan
|
| 19 |
+
language:
|
| 20 |
+
- en
|
| 21 |
+
base_model_relation: quantized
|
| 22 |
+
---
|
| 23 |
+
|
| 24 |
+
<div style="border:2px solid currentColor; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;">
|
| 25 |
+
<div style="border-bottom:1px solid currentColor; padding:6px 12px; font-size:11px; letter-spacing:3px; text-transform:uppercase; opacity:0.7; text-align:center;">PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO · gfx1151</div>
|
| 26 |
+
<div style="padding:14px; display:flex; flex-wrap:wrap; align-items:center; justify-content:center; gap:18px;">
|
| 27 |
+
<pre style="margin:0; flex:0 0 auto; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,monospace; font-size:5px; line-height:1.1; letter-spacing:0;">
|
| 28 |
+
▗▇▇▇▇▇▇▇▖
|
| 29 |
+
▗█▘▝██████▖
|
| 30 |
+
▗▛ ▝██████▆▆▆▆▆▆▆▆▆▆▅
|
| 31 |
+
▟▛ ▗█████████████████▙▖
|
| 32 |
+
▄▄▄▄▄▟▛ ▟████████████████████▖
|
| 33 |
+
▗██▌ ▚▖ ▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔█▘
|
| 34 |
+
▗████▖ ▜▖ ▗█▘
|
| 35 |
+
▜█████▙ ▜▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▀▀▀▀▀▜▙
|
| 36 |
+
▜█████▙ ▝████████████▛ ▜▙
|
| 37 |
+
▜█████▙ ▝██████████▛ ▃ ▜▙
|
| 38 |
+
▀█████▙▖ ▝████████▘ ▟█▙ ▀▙
|
| 39 |
+
▝██████▖ ▝▜█████▘ ▟███▙▂▂▂▂▐█
|
| 40 |
+
▟███████▖ ▜███▘ ▗███████████▛
|
| 41 |
+
▟█████████▄ ▜▛ ▗███████████▀
|
| 42 |
+
▝█████▀ ▗▛ ▗██████▀▀▀▀▀▘
|
| 43 |
+
▜██▘ ▗▛ ▟█████▛▘
|
| 44 |
+
▜█▇▇▇▇▇▇▇▇▇█▖ ▟█████▛
|
| 45 |
+
▝█▖ ▟█████▛
|
| 46 |
+
▝███████▀
|
| 47 |
+
</pre>
|
| 48 |
+
<div style="flex:0 1 auto; max-width:100%; text-align:center;">
|
| 49 |
+
<div style="font-size:23px; font-weight:800; letter-spacing:1px;">FASTCONTEXT-1.0-4B</div>
|
| 50 |
+
<div style="font-size:12.5px; letter-spacing:1px; opacity:0.8; margin-top:5px;"><span style="white-space:nowrap;">4-BIT ROCmFP4</span> · <span style="white-space:nowrap;">QWEN3 DENSE 4B</span> · <span style="white-space:nowrap;">REPO-EXPLORATION SUBAGENT</span> · <span style="white-space:nowrap;">CODE-WEIGHTED IMATRIX</span> · <span style="white-space:nowrap;">SINGLE AMD APU</span></div>
|
| 51 |
+
</div>
|
| 52 |
+
</div>
|
| 53 |
+
<table style="display:table; table-layout:fixed; width:100%; margin:0; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">
|
| 54 |
+
<tr>
|
| 55 |
+
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">FORMAT</div><div style="font-weight:700;">ROCmFP4 4-BIT</div></td>
|
| 56 |
+
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PRECISION</div><div style="font-weight:700;">~4.5 BPW</div></td>
|
| 57 |
+
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">ARCH</div><div style="font-weight:700;">QWEN3 DENSE</div></td>
|
| 58 |
+
<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CONTEXT</div><div style="font-weight:700;">256 K</div></td>
|
| 59 |
+
</tr>
|
| 60 |
+
<tr>
|
| 61 |
+
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PARAMS</div><div style="font-weight:700;">4B DENSE</div></td>
|
| 62 |
+
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">DRAFT</div><div style="font-weight:700;">NO MTP</div></td>
|
| 63 |
+
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">BACKEND</div><div style="font-weight:700;">VULKAN0</div></td>
|
| 64 |
+
<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">LICENSE</div><div style="font-weight:700;">MIT</div></td>
|
| 65 |
+
</tr>
|
| 66 |
+
</table>
|
| 67 |
+
</div>
|
| 68 |
+
|
| 69 |
+
<div style="border:2px solid #dc2626; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;">
|
| 70 |
+
<b style="color:#dc2626; letter-spacing:1px;">⚠ REQUIRES THE ROCmFP4 FORK</b><br>
|
| 71 |
+
The custom <code>q4_0_rocmfp4</code> / <code>q4_0_rocmfp4_fast</code> tensor types <b>will not load in stock llama.cpp, LM Studio, or Ollama</b>. Build/run with <a href="https://github.com/charlie12345/rocmfp4-llama">charlie12345/rocmfp4-llama</a> · branch <code>mtp-rocmfp4-strix</code>.
|
| 72 |
+
</div>
|
| 73 |
+
|
| 74 |
+
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;">
|
| 75 |
+
<b>NOTE //</b> Ignore HuggingFace's auto-detected "F16"/16-bit badge — its parser can't read ROCmFP4 and mislabels the file. These are <b>~4.5 bpw 4-bit</b> ROCmFP4 files; pick by filename in <i>Files and versions</i>.
|
| 76 |
+
</div>
|
| 77 |
+
|
| 78 |
+
Experimental **AMD Strix Halo (gfx1151)** quant of [**microsoft/FastContext-1.0-4B-SFT**](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT) — Microsoft's **repository-exploration subagent** for coding agents. Instead of one model both exploring the repo and solving the task, FastContext is invoked on demand by a main agent, fires **parallel read-only tool calls** (READ / GLOB / GREP), and returns **compact file paths + line ranges** as focused context. Architecturally it's a plain **Qwen3 dense 4B** (`Qwen3ForCausalLM`, 36 layers, hidden 2560, 256K context, MIT-licensed), here in the custom **ROCmFP4** 4-bit format, **imatrix-quantized**.
|
| 79 |
+
|
| 80 |
+
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">01</span> · FILES</div>
|
| 81 |
+
|
| 82 |
+
<div style="overflow:hidden; border-radius:0;">
|
| 83 |
+
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
|
| 84 |
+
<thead><tr>
|
| 85 |
+
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">File</th>
|
| 86 |
+
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Body</th>
|
| 87 |
+
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Size</th>
|
| 88 |
+
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th>
|
| 89 |
+
</tr></thead>
|
| 90 |
+
<tbody>
|
| 91 |
+
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-COHERENT-embF16.gguf</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">all-dual</td><td style="border:1px solid currentColor; padding:7px 10px;">2.8 GB</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>recommended</b> — lowest measured KL vs BF16 (§04)</td></tr>
|
| 92 |
+
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-embF16-imatrix.gguf</code></td><td style="border:1px solid currentColor; padding:7px 10px;">fast</td><td style="border:1px solid currentColor; padding:7px 10px;">2.7 GB</td><td style="border:1px solid currentColor; padding:7px 10px;">~same fidelity, slightly smaller/faster</td></tr>
|
| 93 |
+
</tbody>
|
| 94 |
+
</table>
|
| 95 |
+
</div>
|
| 96 |
+
|
| 97 |
+
Both share genuine **f16 embeddings** (from BF16) + the code-weighted imatrix (see §04). The **COHERENT** build (★) puts every body tensor on the **dual-scale** `q4_0_rocmfp4` kernel — lowest measured KL vs the BF16 reference at ~the same decode speed — vs the STRIX build's faster single-scale `q4_0_rocmfp4_fast` bulk. The Qwen (ChatML) chat template is **baked into the GGUF** — just pass `--jinja`.
|
| 98 |
+
|
| 99 |
+
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
|
| 100 |
+
<b>NOTE // TIED EMBEDDINGS.</b> FastContext has <code>tie_word_embeddings=True</code>, so there's <b>no separate output head</b> — the token-embedding tensor doubles as the lm-head. Setting <code>--token-embedding-type f16</code> therefore gives an <b>f16 embedding <i>and</i> f16 output head</b> in one (no <code>headQ6</code> variant needed — f16 already beats Q6 there).
|
| 101 |
+
</div>
|
| 102 |
+
|
| 103 |
+
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">02</span> · QUICK START</div>
|
| 104 |
+
|
| 105 |
+
Run from the folder holding the `.gguf` (the Qwen ChatML template is baked in — just pass `--jinja`):
|
| 106 |
+
|
| 107 |
+
```bash
|
| 108 |
+
env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
|
| 109 |
+
llama-server \
|
| 110 |
+
-m FastContext-1.0-4B-SFT-ROCmFP4-COHERENT-embF16.gguf \
|
| 111 |
+
--alias fastcontext-4b \
|
| 112 |
+
--host 0.0.0.0 \
|
| 113 |
+
--port 8080 \
|
| 114 |
+
-c 262144 \
|
| 115 |
+
-ctk f16 \
|
| 116 |
+
-ctv f16 \
|
| 117 |
+
--temp 0.7 \
|
| 118 |
+
--top-p 0.8 \
|
| 119 |
+
--top-k 20 \
|
| 120 |
+
-dev Vulkan0 \
|
| 121 |
+
-ngl 999 \
|
| 122 |
+
-fa on \
|
| 123 |
+
-b 2048 \
|
| 124 |
+
-ub 256 \
|
| 125 |
+
-t 16 \
|
| 126 |
+
-tb 16 \
|
| 127 |
+
-cpent 256 \
|
| 128 |
+
-ctxcp 32 \
|
| 129 |
+
--cache-reuse 256 \
|
| 130 |
+
--cache-ram 65536 \
|
| 131 |
+
--jinja \
|
| 132 |
+
--parallel 1 \
|
| 133 |
+
--metrics \
|
| 134 |
+
--no-mmap
|
| 135 |
+
```
|
| 136 |
+
|
| 137 |
+
<div style="overflow:hidden; border-radius:0;">
|
| 138 |
+
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">
|
| 139 |
+
<thead><tr>
|
| 140 |
+
<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px; width:40%;">Flag</th>
|
| 141 |
+
<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Function</th>
|
| 142 |
+
</tr></thead>
|
| 143 |
+
<tbody>
|
| 144 |
+
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>HSA_OVERRIDE_GFX_VERSION=11.5.1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">treat the APU as gfx1151 (Strix Halo)</td></tr>
|
| 145 |
+
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>GGML_HIP_ENABLE_UNIFIED_MEMORY=1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">allow use of the full 128 GB unified memory</td></tr>
|
| 146 |
+
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-dev Vulkan0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">run on Vulkan — fastest backend for ROCmFP4 on Strix Halo</td></tr>
|
| 147 |
+
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ngl 999 · -fa on</code></td><td style="border:1px solid currentColor; padding:6px 10px;">offload all layers · flash attention</td></tr>
|
| 148 |
+
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-c 262144</code></td><td style="border:1px solid currentColor; padding:6px 10px;">context length (256K)</td></tr>
|
| 149 |
+
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-b 2048 · -ub 256 · -t/-tb 16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">prefill batch / micro-batch · CPU threads</td></tr>
|
| 150 |
+
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ctk f16 · -ctv f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">f16 KV cache — how we run it (cheap on a 4B); drop to <code>q8_0</code>/<code>q4_0</code> to use less memory at deep context</td></tr>
|
| 151 |
+
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-cpent · -ctxcp · --cache-reuse · --cache-ram 65536</code></td><td style="border:1px solid currentColor; padding:6px 10px;">cross-turn KV checkpointing + 64 GB resident reuse cache</td></tr>
|
| 152 |
+
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--temp 0.7 --top-p 0.8 --top-k 20</code></td><td style="border:1px solid currentColor; padding:6px 10px;">Qwen3 recommended sampling (instruct/non-thinking)</td></tr>
|
| 153 |
+
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--jinja --parallel 1 --metrics --no-mmap</code></td><td style="border:1px solid currentColor; padding:6px 10px;">apply baked ChatML template · single slot · metrics · weights in RAM</td></tr>
|
| 154 |
+
</tbody>
|
| 155 |
+
</table>
|
| 156 |
+
</div>
|
| 157 |
+
|
| 158 |
+
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
|
| 159 |
+
<b>NOTE //</b> No <code>--spec-*</code> / <code>--spec-type draft-mtp</code> flags — this arch has <b>no MTP head</b> (see §04). It's already fast on its own.
|
| 160 |
+
</div>
|
| 161 |
+
|
| 162 |
+
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">03</span> · USING IT AS A SUBAGENT</div>
|
| 163 |
+
|
| 164 |
+
FastContext isn't a general chat model — it's a **repository-exploration subagent** meant to be **called by your main coding agent**, not driven directly. The intended loop: the main agent delegates "find the relevant context for X" → FastContext issues **parallel read-only tool calls** (`READ`, `GLOB`, `GREP`) → returns **compact file paths + line ranges**, which the main agent folds into its own context to do the actual work. The point is to keep repo-exploration tokens *out* of the main agent's window.
|
| 165 |
+
|
| 166 |
+
- **Chat template:** Qwen (ChatML) is baked into the GGUF — just pass `--jinja`.
|
| 167 |
+
- **Tool calling:** it emits structured `READ`/`GLOB`/`GREP` calls — wire those tools into your harness and use a Qwen/Hermes-style tool-call parser so they're parsed rather than printed. **See the [upstream model card](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT) for the exact subagent protocol + tool schema** (it expects a specific invocation format).
|
| 168 |
+
- **Sampling:** temp `0.7`, top-p `0.8`, top-k `20` (Qwen3 instruct defaults) — already set in §02.
|
| 169 |
+
|
| 170 |
+
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
|
| 171 |
+
<b>NOTE //</b> It's small (4B) and fast (~68 t/s, §04) by design — a cheap, disposable explorer you can fan out in parallel next to a larger main model on the same box. The cross-turn reuse cache (<code>--cache-reuse</code> / <code>--cache-ram</code>) keeps repeated exploration over the same repo cheap.
|
| 172 |
+
</div>
|
| 173 |
+
|
| 174 |
+
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">04</span> · PERFORMANCE & QUALITY</div>
|
| 175 |
+
|
| 176 |
+
<div style="overflow:hidden; border-radius:0;">
|
| 177 |
+
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
|
| 178 |
+
<tbody>
|
| 179 |
+
<tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">DECODE · short context</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~68 t/s (Vulkan / Ryzen AI Max+ 395)</td></tr>
|
| 180 |
+
<tr><td style="border:1px solid currentColor; padding:8px 11px;">SPECULATIVE DECODE</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">none (no MTP head)</td></tr>
|
| 181 |
+
<tr><td style="border:1px solid currentColor; padding:8px 11px;">CONTEXT</td><td style="border:1px solid currentColor; padding:8px 11px;">256K native (dense attention)</td></tr>
|
| 182 |
+
<tr><td style="border:1px solid currentColor; padding:8px 11px;">QUANTIZATION</td><td style="border:1px solid currentColor; padding:8px 11px;">COHERENT body + imatrix (measured win — below)</td></tr>
|
| 183 |
+
</tbody>
|
| 184 |
+
</table>
|
| 185 |
+
</div>
|
| 186 |
+
|
| 187 |
+
**Recommended build = COHERENT (we measured it).** Both builds use f16 tied emb/head + the same imatrix; the lever swept here is the **body kernel**, ranked by **KL divergence vs the true BF16** on held-out code (lower = more faithful). The **all-dual-scale body** (COHERENT) beats the fast-body STRIX build on **every** metric at ~the same decode speed:
|
| 188 |
+
|
| 189 |
+
<div style="overflow:hidden; border-radius:0;">
|
| 190 |
+
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
|
| 191 |
+
<thead><tr>
|
| 192 |
+
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Build (imatrix + embF16, tied head)</th>
|
| 193 |
+
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Body</th>
|
| 194 |
+
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Mean KLD vs BF16 ↓</th>
|
| 195 |
+
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Median KLD ↓</th>
|
| 196 |
+
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Top-token</th>
|
| 197 |
+
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">PPL(Q) ↓</th>
|
| 198 |
+
</tr></thead>
|
| 199 |
+
<tbody>
|
| 200 |
+
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>COHERENT</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">all-dual</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>0.03422</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>0.00955</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>92.08%</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>4.192</b></td></tr>
|
| 201 |
+
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>STRIX</code></td><td style="border:1px solid currentColor; padding:7px 10px;">fast</td><td style="border:1px solid currentColor; padding:7px 10px;">0.03934</td><td style="border:1px solid currentColor; padding:7px 10px;">0.01016</td><td style="border:1px solid currentColor; padding:7px 10px;">91.38%</td><td style="border:1px solid currentColor; padding:7px 10px;">4.213</td></tr>
|
| 202 |
+
</tbody>
|
| 203 |
+
</table>
|
| 204 |
+
</div>
|
| 205 |
+
|
| 206 |
+
A **clean sweep**: COHERENT is lower on mean KLD (−13%), median KLD (−6%), RMS Δp (6.43% vs 6.95%), **and** perplexity (4.192 vs 4.213), and higher on same-top-token (+0.70 pp) — every metric, same direction (BF16 reference PPL 4.074). So it's the default; STRIX stays as a marginally smaller/faster fallback.
|
| 207 |
+
|
| 208 |
+
**Fast on its own.** ~68 t/s short-context decode on a Ryzen AI Max+ 395 (Vulkan0, measured `llama-bench tg128`). It's a 4B dense Qwen3 with **no MTP head**, so there's no speculative decoding — it doesn't need it, and at 4B it's a cheap explorer you can run several of in parallel.
|
| 209 |
+
|
| 210 |
+
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
|
| 211 |
+
<b>NOTE // imatrix.</b> Both builds are quantized <b>with</b> an importance matrix (Kalomaze <code>groups_merged</code> + froggeric <code>code</code>/<code>technical</code>, via <a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a>), computed on this model's BF16. We measured the <b>COHERENT-vs-STRIX</b> comparison above (both imatrix); we did <b>not</b> run a separate imatrix-vs-no-imatrix ablation on this model. Scope: the KL/PPL figures are a fidelity-vs-BF16 measurement on a held-out code slice, <b>not</b> an absolute coding benchmark.
|
| 212 |
+
</div>
|
| 213 |
+
|
| 214 |
+
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> · BUILD (REPRODUCIBLE)</div>
|
| 215 |
+
|
| 216 |
+
```bash
|
| 217 |
+
# 0) convert the safetensors -> BF16 GGUF (plain qwen3 dense; no MTP, tied embeddings)
|
| 218 |
+
python convert_hf_to_gguf.py FastContext-1.0-4B-SFT/ --outtype bf16 --outfile FastContext-1.0-4B-SFT-BF16.gguf
|
| 219 |
+
|
| 220 |
+
# 1) imatrix on the BF16 (general+code: Kalomaze groups_merged + froggeric code/technical)
|
| 221 |
+
llama-imatrix -m FastContext-1.0-4B-SFT-BF16.gguf -f general+code-calib.txt -o fastcontext-4b.imatrix -c 512 -ngl 999
|
| 222 |
+
|
| 223 |
+
# 2) RECOMMENDED: COHERENT all-dual body + f16 tied emb/head (the ★ file) — lowest KL (§04).
|
| 224 |
+
# tie_word_embeddings=True -> --token-embedding-type f16 also gives an f16 output head; no --output-tensor-type.
|
| 225 |
+
llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
|
| 226 |
+
FastContext-1.0-4B-SFT-BF16.gguf FastContext-1.0-4B-SFT-ROCmFP4-COHERENT-embF16.gguf Q4_0_ROCMFP4_COHERENT
|
| 227 |
+
|
| 228 |
+
# fast-body STRIX fallback (same f16 emb + imatrix)
|
| 229 |
+
llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
|
| 230 |
+
FastContext-1.0-4B-SFT-BF16.gguf FastContext-1.0-4B-SFT-ROCmFP4-STRIX-embF16-imatrix.gguf Q4_0_ROCMFP4_STRIX
|
| 231 |
+
```
|
| 232 |
+
|
| 233 |
+
> Experimental research build for AMD Strix Halo — hardware/driver/prompt-sensitive, may not reproduce elsewhere. Not native FP4 tensor-core execution.
|
| 234 |
+
|
| 235 |
+
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">06</span> · LINEAGE & CREDITS</div>
|
| 236 |
+
|
| 237 |
+
<div style="overflow:hidden; border-radius:0;">
|
| 238 |
+
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
|
| 239 |
+
<tbody>
|
| 240 |
+
<tr><td style="border:1px solid currentColor; padding:8px 11px; width:26%;">BASE MODEL</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/microsoft/FastContext-1.0-4B-SFT">microsoft/FastContext-1.0-4B-SFT</a> (MIT, Microsoft) · repository-exploration subagent · Qwen3 dense 4B (<code>Qwen3ForCausalLM</code>)</td></tr>
|
| 241 |
+
<tr><td style="border:1px solid currentColor; padding:8px 11px;">CALIBRATION</td><td style="border:1px solid currentColor; padding:8px 11px;">Kalomaze <code>groups_merged</code> + froggeric <code>code</code>/<code>technical</code> via <a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a></td></tr>
|
| 242 |
+
<tr><td style="border:1px solid currentColor; padding:8px 11px;">FORMAT + RUNTIME</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://github.com/charlie12345/rocmfp4-llama">charlie12345/rocmfp4-llama</a> (based on llama.cpp, MIT)</td></tr>
|
| 243 |
+
</tbody>
|
| 244 |
+
</table>
|
| 245 |
+
</div>
|
| 246 |
+
|
| 247 |
+
*Derivative quantization — verify the base model's license before redistribution / use.*
|