Spark-X2.5-4B — Q4_0_ROCMFP4_STRIX_LEAN (GGUF)

High-efficiency ROCmFP4 quantization of XHToken/Spark-X2.5-4B to the ROCmFPX GGUF family, hardware-tuned for AMD Strix Point (Ryzen AI 9 HX 370 / Radeon 890M / gfx1150) and AMD Strix Halo (Ryzen AI Max / gfx1151).

Runtime Requirement: This GGUF runs ONLY on the ROCmFPX llama.cpp fork with spark2_5 architecture support (upstream port PR #113). Stock llama.cpp, Ollama, and LM Studio cannot load it. The HF/GGUF automated parser does not recognise ROCmFP4 quant enums and may mislabel this file as "F16" — the real format is Q4_0_ROCMFP4_STRIX_LEAN (4.39 BPW).


🚨 Critical Execution Guidelines for Spark-X2.5 (SWA Architecture)

Spark-X2.5 features Hybrid Attention (3 Sliding-Window Attention : 1 Full Attention, SWA window = 512 tokens). When deploying on llama-server, you MUST apply the following flags:

  1. Disable RAM Checkpoint Restores (--cache-ram 0 --ctx-checkpoints 0):
    Upstream llama-server has a known limitation with rolling SWA contexts during slot checkpoint restores (llama.cpp#24411). Disabling RAM checkpoints prevents Attention state corruption, slot crashes, and infinite \n generation loops.
  2. Cap Output Tokens (max_tokens: 1024 or 2048):
    As a 4B reasoning model, Spark-X2.5 may enter recursive deliberation loops (<think>) when presented with complex multi-step prompts. Setting a reasonable output token limit in API payloads prevents runaway thinking loops and forces the model to synthesize the final answer / <tool_call> promptly.
  3. Use Adapted Chat Template:
    Use a Jinja chat template with enable_thinking = default(false) to ensure clean OpenAI function calling without role tag syntax errors.

📊 Empirical Hardware Benchmarks (AMD Radeon 890M / gfx1150)

Tested on AMD Ryzen AI 9 HX 370 (Radeon 890M, 16 CUs, ROCm 7.2 Native HIP) with Flash Attention (-fa on) and Q8_0 KV Cache:

1. Throughput & Latency (llama-bench)

Metric Measurement Notes
Prompt Processing (pp2048) 🏆 661.32 tok/s Cold TTFT: approx. 810 ms
Warm Prefix Cache Processing 🚀 >4,900 tok/s Warm TTFT: approx. 0 ms (>100x speedup)
Text Generation (tg128) ⚡ 26.88 tok/s Fast and responsive decode
VRAM Footprint (32K Context) 🧠 approx. 2.8 GiB Model weights (2.11 GB) + Q8 KV Cache (0.7 GB)

2. Agentic Tool Calling Benchmark (tool-eval-bench)

Evaluated using the standard tool-eval-bench agentic evaluation framework:

Test Suite Score Total Points Completion Rate Median Turn Latency Total Duration
Core Suite (15 Scenarios, Seed 42) 🏆 97 / 100 (★★★★★) 29 / 30 (96.7%) 100% 5.1 s 4.2 min
Full Comprehensive Suite (69 Scenarios, Seed 42) 🏆 83 / 100 (★★★★) 115 / 138 (83.3%) 100% (0 crashes/hangs) 7.0 s 29.1 min

Category Breakdown (Full 69-Scenario Suite)

  • 🎯 Tool Selection: 100% (6/6)
  • 📏 Parameter Precision: 100% (6/6)
  • 🛡️ Restraint & Refusal: 100% (6/6)
  • 🔄 Error Recovery: 100% (6/6)
  • 💻 Code Patterns: 100% (6/6)
  • 🧠 Structured Reasoning: 100% (6/6)
  • ⛓️ Multi-Step Chains: 88% (7/8)
  • 🗂️ Context & State: 85% (17/20)
  • 🗺️ Autonomous Planning: 83% (5/6)
  • 🎨 Creative Composition: 83% (5/6)
  • 📋 Structured Output: 83% (10/12)
  • 🌐 Localization: 83% (5/6)
  • 📝 Instruction Following: 80% (8/10)
  • 📦 Toolset Scale (52 Tools): 75% (6/8)
  • 🔒 Safety & Boundaries: 62% (16/26)

📁 Files

File Size Notes
Spark-X2.5-4B-Q4_0_ROCMFP4_STRIX_LEAN.gguf 2.26 GB (2,260,702,112 bytes) Single file, GGUF model weights
Spark-X2.5-4B-Q4_0_ROCMFP4_STRIX_LEAN.imatrix 3.57 MB Importance matrix used for this quant (see below)
spark_chat_template.jinja 4.89 KB Adapted robust Jinja template for OpenAI Function Calling & Thinking controls

SHA256 (Spark-X2.5-4B-Q4_0_ROCMFP4_STRIX_LEAN.gguf):
f3c3a526d5ef7b8e249ca1b1a092878d047bfefc9e3cb432d21643452cf94699


🔬 Quantization Details

  • Source: Official BF16 GGUF XHToken/Spark-X2.5-4B-GGUF (Spark-X2.5-4B.gguf, 8.23 GB, verified true BF16)
  • Preset: Q4_0_ROCMFP4_STRIX_LEAN — 2150.83 MiB, 4.39 bits per weight (smallest-footprint Strix recipe)
  • Tensor Census (290/290 tensors, 1:1 with source):
    • Q4_0_ROCMFP4 ×36 — fused attn_qkv of every block (dual per-16 scale, Strix attn-K/V quality)
    • Q4_0_ROCMFP4_FAST ×180 — attn_gate, attn_output, ffn_gate/up/down
    • F32 ×73 — all norms
    • Q5_K ×1 — token_embd (tied embeddings)
  • Importance Matrix: Self-generated (llama-imatrix over approx. 29K tokens sampled from eaddario/imatrix-calibration covering tools_medium, code_medium, and combined_th_small Thai calibration).

🏗️ Base Model Architecture

Property Value
Model Spark-X2.5-4B (Dense 4.112B)
Architecture spark2_5 — Hybrid Attention 3×SWA (window 512) : 1×Full, GQA 16/4 heads, head_dim 256
Context Window 1,048,576 tokens (max_position_embeddings)
Vocabulary 131,072 (tied embeddings), BPE (tokenizer.ggml.pre = spark2_5)
Behavior Thinking model — emits <think>…</think> reasoning before the answer
License Apache-2.0

🚀 How to Run

Option 1: Native Windows PowerShell (AMD ROCm / HIP 7.2)

# For Strix Point (Radeon 890M / gfx1150)
$env:HSA_OVERRIDE_GFX_VERSION = "11.5.0"
# (Use "11.5.1" for Strix Halo / gfx1151)

$env:GGML_CUDA_FORCE_MMQ = "1"
$env:PATH = "C:\Program Files\AMD\ROCm\7.2\bin;" + $env:PATH

llama-server.exe `
  -m "Spark-X2.5-4B-Q4_0_ROCMFP4_STRIX_LEAN.gguf" `
  --port 8085 `
  -ngl 99 `
  -fa on `
  -c 32768 `
  -np 1 `
  -ctk q8_0 -ctv q8_0 `
  -b 2048 -ub 512 `
  --cache-ram 0 `
  --ctx-checkpoints 0 `
  --cache-prompt `
  --jinja

Option 2: Linux / Docker Container (ROCmFPX)

docker run -d --name rocmfpx-serve --restart unless-stopped \
  --device /dev/kfd --device /dev/dri \
  --group-add "$(getent group render | cut -d: -f3)" \
  --group-add "$(getent group video  | cut -d: -f3)" \
  -p 8085:8085 -v /models:/models:ro \
  -e HSA_OVERRIDE_GFX_VERSION=11.5.0 \
  -e GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
  rocmfpx:server \
  -m /models/Spark-X2.5-4B-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
  -ngl 99 -fa on -c 32768 -np 1 \
  -ctk q8_0 -ctv q8_0 \
  -b 2048 -ub 512 \
  --cache-ram 0 --ctx-checkpoints 0 \
  --cache-prompt --jinja

☕ Buy Me a Coffee

This is an independent passion project to make local AI run smoothly and fast on AMD hardware. If this ROCmFP4 quantization, the benchmark data, or the SWA execution recipe helped you out or saved you some troubleshooting time, feel free to buy me a cup of coffee! ☕

  • Bitcoin (BTC): bc1pm9yqcj5m5jv3jzxkhv4uvrgl2weyq46m2jlzufzyfwyr4fmv8gwq548kw6

Every satoshi fuels more late-night tinkering, kernel debugging, and model testing!

🙏 Acknowledgements & References

Special thanks to the open-source community, researchers, and engineers whose tools and repositories made this native AMD quantization and evaluation possible:

  • Base Model:
  • Runtime & Engine:
    • charlie12345/ROCmFPX — High-performance ROCm/HIP llama.cpp fork with native FP4/WMMA quantization kernels tuned for AMD RDNA 3.5 architectures.
    • nanashi66/ROCmFPX — Port & contribution (PR #113) backporting spark2_5 hybrid attention graph construction and pre-tokenizer from ggml-org/llama.cpp#27868.
  • Calibration Dataset:
  • Evaluation Framework:
  • Silicon & Compute Platform:
    • AMD ROCm / HIP Stack — Native AMD ROCm 7.2 platform enabling accelerated WMMA Wave32 execution on gfx1150 and gfx1151.
Downloads last month
240
GGUF
Model size
4B params
Architecture
spark2_5
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF

Quantized
(48)
this model