plunderstruck's picture
Upload README.md with huggingface_hub
2fd4bb5 verified
|
Raw History Blame
23.8 kB
metadata
base_model: microsoft/FastContext-1.0-4B-SFT
license: mit
library_name: gguf
tags:
  - gguf
  - rocmfp4
  - qwen3
  - fastcontext
  - subagent
  - repository-exploration
  - coder
  - agentic
  - imatrix
  - strix-halo
  - amd
  - rocm
  - vulkan
language:
  - en
base_model_relation: quantized
PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO Β· gfx1151
            β–—β–‡β–‡β–‡β–‡β–‡β–‡β–‡β––                 
           β–—β–ˆβ–˜β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ––                
          β–—β–›   β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–†β–†β–†β–†β–†β–†β–†β–†β–†β–†β–…     
         β–Ÿβ–›    β–—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–™β––   
   β–„β–„β–„β–„β–„β–Ÿβ–›    β–Ÿβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ––  
 β–—β–ˆβ–ˆβ–Œ    β–šβ––   β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–”β–ˆβ–˜  
β–—β–ˆβ–ˆβ–ˆβ–ˆβ––    β–œβ––                    β–—β–ˆβ–˜   
β–œβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–™    β–œβ–†β–†β–†β–†β–†β–†β–†β–†β–†β–†β–†β–†β–†β–†β–†β–€β–€β–€β–€β–€β–œβ–™    
 β–œβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–™    β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–›       β–œβ–™   
  β–œβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–™    β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–›    β–ƒ    β–œβ–™  
   β–€β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–™β––   β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–˜    β–Ÿβ–ˆβ–™    β–€β–™ 
    β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ––   β–β–œβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–˜    β–Ÿβ–ˆβ–ˆβ–ˆβ–™β–‚β–‚β–‚β–‚β–β–ˆ
    β–Ÿβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ––    β–œβ–ˆβ–ˆβ–ˆβ–˜   β–—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–›
   β–Ÿβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–„    β–œβ–›    β–—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–€ 
  β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–€        β–—β–›    β–—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–€β–€β–€β–€β–€β–˜  
    β–œβ–ˆβ–ˆβ–˜        β–—β–›    β–Ÿβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–›β–˜        
     β–œβ–ˆβ–‡β–‡β–‡β–‡β–‡β–‡β–‡β–‡β–‡β–ˆβ––   β–Ÿβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–›          
                β–β–ˆβ–– β–Ÿβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–›           
                 β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–€            
FASTCONTEXT-1.0-4B
4-BIT ROCmFP4 Β· QWEN3 DENSE 4B Β· REPO-EXPLORATION SUBAGENT Β· CODE-WEIGHTED IMATRIX Β· SINGLE AMD APU
FORMAT
ROCmFP4 4-BIT
PRECISION
~4.5 BPW
ARCH
QWEN3 DENSE
CONTEXT
256 K
PARAMS
4B DENSE
DRAFT
NO MTP
BACKEND
VULKAN0
LICENSE
MIT
⚠ REQUIRES THE ROCmFP4 FORK
The custom q4_0_rocmfp4 / q4_0_rocmfp4_fast tensor types will not load in stock llama.cpp, LM Studio, or Ollama. Build/run with charlie12345/rocmfp4-llama Β· branch mtp-rocmfp4-strix.
NOTE // Ignore HuggingFace's auto-detected "F16"/16-bit badge β€” its parser can't read ROCmFP4 and mislabels the file. These are ~4.5 bpw 4-bit ROCmFP4 files; pick by filename in Files and versions.

Experimental AMD Strix Halo (gfx1151) quant of microsoft/FastContext-1.0-4B-SFT β€” Microsoft's repository-exploration subagent for coding agents. Instead of one model both exploring the repo and solving the task, FastContext is invoked on demand by a main agent, fires parallel read-only tool calls (READ / GLOB / GREP), and returns compact file paths + line ranges as focused context. Architecturally it's a plain Qwen3 dense 4B (Qwen3ForCausalLM, 36 layers, hidden 2560, 256K context, MIT-licensed), here in the custom ROCmFP4 4-bit format, imatrix-quantized.

01 Β· FILES
File Body Size Pick if
…-COHERENT-embF16.gguf β˜…all-dual2.8 GBrecommended β€” lowest measured KL vs BF16 (Β§04)
…-STRIX-embF16-imatrix.gguffast2.7 GB~same fidelity, slightly smaller/faster

Both share genuine f16 embeddings (from BF16) + the code-weighted imatrix (see Β§04). The COHERENT build (β˜…) puts every body tensor on the dual-scale q4_0_rocmfp4 kernel β€” lowest measured KL vs the BF16 reference at ~the same decode speed β€” vs the STRIX build's faster single-scale q4_0_rocmfp4_fast bulk. The Qwen (ChatML) chat template is baked into the GGUF β€” just pass --jinja.

NOTE // TIED EMBEDDINGS. FastContext has tie_word_embeddings=True, so there's no separate output head β€” the token-embedding tensor doubles as the lm-head. Setting --token-embedding-type f16 therefore gives an f16 embedding and f16 output head in one (no headQ6 variant needed β€” f16 already beats Q6 there).
02 Β· QUICK START

Run from the folder holding the .gguf (the Qwen ChatML template is baked in β€” just pass --jinja):

env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server \
  -m FastContext-1.0-4B-SFT-ROCmFP4-COHERENT-embF16.gguf \
  --alias fastcontext-4b \
  --host 0.0.0.0 \
  --port 8080 \
  -c 262144 \
  -ctk f16 \
  -ctv f16 \
  --temp 0.7 \
  --top-p 0.8 \
  --top-k 20 \
  -dev Vulkan0 \
  -ngl 999 \
  -fa on \
  -b 2048 \
  -ub 256 \
  -t 16 \
  -tb 16 \
  -cpent 256 \
  -ctxcp 32 \
  --cache-reuse 256 \
  --cache-ram 65536 \
  --jinja \
  --parallel 1 \
  --metrics \
  --no-mmap
Flag Function
HSA_OVERRIDE_GFX_VERSION=11.5.1treat the APU as gfx1151 (Strix Halo)
GGML_HIP_ENABLE_UNIFIED_MEMORY=1allow use of the full 128 GB unified memory
-dev Vulkan0run on Vulkan β€” fastest backend for ROCmFP4 on Strix Halo
-ngl 999 Β· -fa onoffload all layers Β· flash attention
-c 262144context length (256K)
-b 2048 Β· -ub 256 Β· -t/-tb 16prefill batch / micro-batch Β· CPU threads
-ctk f16 Β· -ctv f16f16 KV cache β€” how we run it (cheap on a 4B); drop to q8_0/q4_0 to use less memory at deep context
-cpent Β· -ctxcp Β· --cache-reuse Β· --cache-ram 65536cross-turn KV checkpointing + 64 GB resident reuse cache
--temp 0.7 --top-p 0.8 --top-k 20Qwen3 recommended sampling (instruct/non-thinking)
--jinja --parallel 1 --metrics --no-mmapapply baked ChatML template Β· single slot Β· metrics Β· weights in RAM
NOTE // No --spec-* / --spec-type draft-mtp flags β€” this arch has no MTP head (see Β§04). It's already fast on its own.
03 Β· USING IT AS A SUBAGENT

FastContext isn't a general chat model β€” it's a repository-exploration subagent meant to be called by your main coding agent, not driven directly. The intended loop: the main agent delegates "find the relevant context for X" β†’ FastContext issues parallel read-only tool calls (READ, GLOB, GREP) β†’ returns compact file paths + line ranges, which the main agent folds into its own context to do the actual work. The point is to keep repo-exploration tokens out of the main agent's window.

  • Chat template: Qwen (ChatML) is baked into the GGUF β€” just pass --jinja.
  • Tool calling: it emits structured READ/GLOB/GREP calls β€” wire those tools into your harness and use a Qwen/Hermes-style tool-call parser so they're parsed rather than printed. See the upstream model card for the exact subagent protocol + tool schema (it expects a specific invocation format).
  • Sampling: temp 0.7, top-p 0.8, top-k 20 (Qwen3 instruct defaults) β€” already set in Β§02.
NOTE // It's small (4B) and fast (~68 t/s, Β§04) by design β€” a cheap, disposable explorer you can fan out in parallel next to a larger main model on the same box. The cross-turn reuse cache (--cache-reuse / --cache-ram) keeps repeated exploration over the same repo cheap.
04 Β· PERFORMANCE & QUALITY
DECODE Β· short context~68 t/s (Vulkan / Ryzen AI Max+ 395)
SPECULATIVE DECODEnone (no MTP head)
CONTEXT256K native (dense attention)
QUANTIZATIONCOHERENT body + imatrix (measured win β€” below)

Recommended build = COHERENT (we measured it). Both builds use f16 tied emb/head + the same imatrix; the lever swept here is the body kernel, ranked by KL divergence vs the true BF16 on held-out code (lower = more faithful). The all-dual-scale body (COHERENT) beats the fast-body STRIX build on every metric at ~the same decode speed:

Build (imatrix + embF16, tied head) Body Mean KLD vs BF16 ↓ Median KLD ↓ Top-token PPL(Q) ↓
COHERENT β˜…all-dual0.034220.0095592.08%4.192
STRIXfast0.039340.0101691.38%4.213

A clean sweep: COHERENT is lower on mean KLD (βˆ’13%), median KLD (βˆ’6%), RMS Ξ”p (6.43% vs 6.95%), and perplexity (4.192 vs 4.213), and higher on same-top-token (+0.70 pp) β€” every metric, same direction (BF16 reference PPL 4.074). So it's the default; STRIX stays as a marginally smaller/faster fallback.

Fast on its own. ~68 t/s short-context decode on a Ryzen AI Max+ 395 (Vulkan0, measured llama-bench tg128). It's a 4B dense Qwen3 with no MTP head, so there's no speculative decoding β€” it doesn't need it, and at 4B it's a cheap explorer you can run several of in parallel.

NOTE // imatrix. Both builds are quantized with an importance matrix (Kalomaze groups_merged + froggeric code/technical, via froggeric/imatrix), computed on this model's BF16. We measured the COHERENT-vs-STRIX comparison above (both imatrix); we did not run a separate imatrix-vs-no-imatrix ablation on this model. Scope: the KL/PPL figures are a fidelity-vs-BF16 measurement on a held-out code slice, not an absolute coding benchmark.
05 Β· BUILD (REPRODUCIBLE)
# 0) convert the safetensors -> BF16 GGUF (plain qwen3 dense; no MTP, tied embeddings)
python convert_hf_to_gguf.py FastContext-1.0-4B-SFT/ --outtype bf16 --outfile FastContext-1.0-4B-SFT-BF16.gguf

# 1) imatrix on the BF16 (general+code: Kalomaze groups_merged + froggeric code/technical)
llama-imatrix -m FastContext-1.0-4B-SFT-BF16.gguf -f general+code-calib.txt -o fastcontext-4b.imatrix -c 512 -ngl 999

# 2) RECOMMENDED: COHERENT all-dual body + f16 tied emb/head (the β˜… file) β€” lowest KL (Β§04).
#    tie_word_embeddings=True -> --token-embedding-type f16 also gives an f16 output head; no --output-tensor-type.
llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
  FastContext-1.0-4B-SFT-BF16.gguf  FastContext-1.0-4B-SFT-ROCmFP4-COHERENT-embF16.gguf  Q4_0_ROCMFP4_COHERENT

# fast-body STRIX fallback (same f16 emb + imatrix)
llama-quantize --token-embedding-type f16 --imatrix fastcontext-4b.imatrix \
  FastContext-1.0-4B-SFT-BF16.gguf  FastContext-1.0-4B-SFT-ROCmFP4-STRIX-embF16-imatrix.gguf  Q4_0_ROCMFP4_STRIX

Experimental research build for AMD Strix Halo β€” hardware/driver/prompt-sensitive, may not reproduce elsewhere. Not native FP4 tensor-core execution.

06 Β· LINEAGE & CREDITS
BASE MODELmicrosoft/FastContext-1.0-4B-SFT (MIT, Microsoft) Β· repository-exploration subagent Β· Qwen3 dense 4B (Qwen3ForCausalLM)
CALIBRATIONKalomaze groups_merged + froggeric code/technical via froggeric/imatrix
FORMAT + RUNTIMEcharlie12345/rocmfp4-llama (based on llama.cpp, MIT)

Derivative quantization β€” verify the base model's license before redistribution / use.