Instructions to use nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: llama cli -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: llama cli -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: ./llama-cli -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: ./build/bin/llama-cli -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
Use Docker
docker model run hf.co/nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
- LM Studio
- Jan
- vLLM
How to use nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
- Ollama
How to use nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF with Ollama:
ollama run hf.co/nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
- Unsloth Desktop
- Pi
How to use nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF with Docker Model Runner:
docker model run hf.co/nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
- Lemonade
How to use nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
Run and chat with the model
lemonade run user.Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF-Q4_0_ROCMFP
List all available models
lemonade list
- Hermes Agent
How to use nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "nanash66/Spark-X2.5-4B-ROCmFP4-STRIX_LEAN-GGUF:Q4_0_ROCMFP" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Spark-X2.5-4B — Q4_0_ROCMFP4_STRIX_LEAN (GGUF)
High-efficiency ROCmFP4 quantization of XHToken/Spark-X2.5-4B to the ROCmFPX GGUF family, hardware-tuned for AMD Strix Point (Ryzen AI 9 HX 370 / Radeon 890M / gfx1150) and AMD Strix Halo (Ryzen AI Max / gfx1151).
Runtime Requirement: This GGUF runs ONLY on the ROCmFPX llama.cpp fork with
spark2_5architecture support (upstream port PR #113). Stock llama.cpp, Ollama, and LM Studio cannot load it. The HF/GGUF automated parser does not recognise ROCmFP4 quant enums and may mislabel this file as "F16" — the real format isQ4_0_ROCMFP4_STRIX_LEAN(4.39 BPW).
🚨 Critical Execution Guidelines for Spark-X2.5 (SWA Architecture)
Spark-X2.5 features Hybrid Attention (3 Sliding-Window Attention : 1 Full Attention, SWA window = 512 tokens). When deploying on
llama-server, you MUST apply the following flags:
- Disable RAM Checkpoint Restores (
--cache-ram 0 --ctx-checkpoints 0):
Upstreamllama-serverhas a known limitation with rolling SWA contexts during slot checkpoint restores (llama.cpp#24411). Disabling RAM checkpoints prevents Attention state corruption, slot crashes, and infinite\ngeneration loops.- Cap Output Tokens (
max_tokens: 1024or2048):
As a 4B reasoning model, Spark-X2.5 may enter recursive deliberation loops (<think>) when presented with complex multi-step prompts. Setting a reasonable output token limit in API payloads prevents runaway thinking loops and forces the model to synthesize the final answer /<tool_call>promptly.- Use Adapted Chat Template:
Use a Jinja chat template withenable_thinking = default(false)to ensure clean OpenAI function calling without role tag syntax errors.
📊 Empirical Hardware Benchmarks (AMD Radeon 890M / gfx1150)
Tested on AMD Ryzen AI 9 HX 370 (Radeon 890M, 16 CUs, ROCm 7.2 Native HIP) with Flash Attention (-fa on) and Q8_0 KV Cache:
1. Throughput & Latency (llama-bench)
| Metric | Measurement | Notes |
|---|---|---|
Prompt Processing (pp2048) |
🏆 661.32 tok/s |
Cold TTFT: approx. 810 ms |
| Warm Prefix Cache Processing | 🚀 >4,900 tok/s |
Warm TTFT: approx. 0 ms (>100x speedup) |
Text Generation (tg128) |
⚡ 26.88 tok/s |
Fast and responsive decode |
| VRAM Footprint (32K Context) | 🧠 approx. 2.8 GiB | Model weights (2.11 GB) + Q8 KV Cache (0.7 GB) |
2. Agentic Tool Calling Benchmark (tool-eval-bench)
Evaluated using the standard tool-eval-bench agentic evaluation framework:
| Test Suite | Score | Total Points | Completion Rate | Median Turn Latency | Total Duration |
|---|---|---|---|---|---|
| Core Suite (15 Scenarios, Seed 42) | 🏆 97 / 100 (★★★★★) |
29 / 30 (96.7%) | 100% | 5.1 s |
4.2 min |
| Full Comprehensive Suite (69 Scenarios, Seed 42) | 🏆 83 / 100 (★★★★) |
115 / 138 (83.3%) | 100% (0 crashes/hangs) | 7.0 s |
29.1 min |
Category Breakdown (Full 69-Scenario Suite)
- 🎯 Tool Selection:
100%(6/6) - 📏 Parameter Precision:
100%(6/6) - 🛡️ Restraint & Refusal:
100%(6/6) - 🔄 Error Recovery:
100%(6/6) - 💻 Code Patterns:
100%(6/6) - 🧠 Structured Reasoning:
100%(6/6) - ⛓️ Multi-Step Chains:
88%(7/8) - 🗂️ Context & State:
85%(17/20) - 🗺️ Autonomous Planning:
83%(5/6) - 🎨 Creative Composition:
83%(5/6) - 📋 Structured Output:
83%(10/12) - 🌐 Localization:
83%(5/6) - 📝 Instruction Following:
80%(8/10) - 📦 Toolset Scale (52 Tools):
75%(6/8) - 🔒 Safety & Boundaries:
62%(16/26)
📁 Files
| File | Size | Notes |
|---|---|---|
Spark-X2.5-4B-Q4_0_ROCMFP4_STRIX_LEAN.gguf |
2.26 GB (2,260,702,112 bytes) | Single file, GGUF model weights |
Spark-X2.5-4B-Q4_0_ROCMFP4_STRIX_LEAN.imatrix |
3.57 MB | Importance matrix used for this quant (see below) |
spark_chat_template.jinja |
4.89 KB | Adapted robust Jinja template for OpenAI Function Calling & Thinking controls |
SHA256 (Spark-X2.5-4B-Q4_0_ROCMFP4_STRIX_LEAN.gguf):f3c3a526d5ef7b8e249ca1b1a092878d047bfefc9e3cb432d21643452cf94699
🔬 Quantization Details
- Source: Official BF16 GGUF
XHToken/Spark-X2.5-4B-GGUF(Spark-X2.5-4B.gguf, 8.23 GB, verified true BF16) - Preset:
Q4_0_ROCMFP4_STRIX_LEAN— 2150.83 MiB, 4.39 bits per weight (smallest-footprint Strix recipe) - Tensor Census (290/290 tensors, 1:1 with source):
Q4_0_ROCMFP4×36 — fusedattn_qkvof every block (dual per-16 scale, Strix attn-K/V quality)Q4_0_ROCMFP4_FAST×180 —attn_gate,attn_output,ffn_gate/up/downF32×73 — all normsQ5_K×1 —token_embd(tied embeddings)
- Importance Matrix: Self-generated (
llama-imatrixover approx. 29K tokens sampled from eaddario/imatrix-calibration coveringtools_medium,code_medium, andcombined_th_smallThai calibration).
🏗️ Base Model Architecture
| Property | Value |
|---|---|
| Model | Spark-X2.5-4B (Dense 4.112B) |
| Architecture | spark2_5 — Hybrid Attention 3×SWA (window 512) : 1×Full, GQA 16/4 heads, head_dim 256 |
| Context Window | 1,048,576 tokens (max_position_embeddings) |
| Vocabulary | 131,072 (tied embeddings), BPE (tokenizer.ggml.pre = spark2_5) |
| Behavior | Thinking model — emits <think>…</think> reasoning before the answer |
| License | Apache-2.0 |
🚀 How to Run
Option 1: Native Windows PowerShell (AMD ROCm / HIP 7.2)
# For Strix Point (Radeon 890M / gfx1150)
$env:HSA_OVERRIDE_GFX_VERSION = "11.5.0"
# (Use "11.5.1" for Strix Halo / gfx1151)
$env:GGML_CUDA_FORCE_MMQ = "1"
$env:PATH = "C:\Program Files\AMD\ROCm\7.2\bin;" + $env:PATH
llama-server.exe `
-m "Spark-X2.5-4B-Q4_0_ROCMFP4_STRIX_LEAN.gguf" `
--port 8085 `
-ngl 99 `
-fa on `
-c 32768 `
-np 1 `
-ctk q8_0 -ctv q8_0 `
-b 2048 -ub 512 `
--cache-ram 0 `
--ctx-checkpoints 0 `
--cache-prompt `
--jinja
Option 2: Linux / Docker Container (ROCmFPX)
docker run -d --name rocmfpx-serve --restart unless-stopped \
--device /dev/kfd --device /dev/dri \
--group-add "$(getent group render | cut -d: -f3)" \
--group-add "$(getent group video | cut -d: -f3)" \
-p 8085:8085 -v /models:/models:ro \
-e HSA_OVERRIDE_GFX_VERSION=11.5.0 \
-e GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
rocmfpx:server \
-m /models/Spark-X2.5-4B-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
-ngl 99 -fa on -c 32768 -np 1 \
-ctk q8_0 -ctv q8_0 \
-b 2048 -ub 512 \
--cache-ram 0 --ctx-checkpoints 0 \
--cache-prompt --jinja
☕ Buy Me a Coffee
This is an independent passion project to make local AI run smoothly and fast on AMD hardware. If this ROCmFP4 quantization, the benchmark data, or the SWA execution recipe helped you out or saved you some troubleshooting time, feel free to buy me a cup of coffee! ☕
- Bitcoin (BTC):
bc1pm9yqcj5m5jv3jzxkhv4uvrgl2weyq46m2jlzufzyfwyr4fmv8gwq548kw6
Every satoshi fuels more late-night tinkering, kernel debugging, and model testing!
🙏 Acknowledgements & References
Special thanks to the open-source community, researchers, and engineers whose tools and repositories made this native AMD quantization and evaluation possible:
- Base Model:
- XHToken/Spark-X2.5-4B — Base model weights and architecture by XHToken (Apache-2.0).
XHToken/Spark-X2.5-4B-GGUF— Official baseline BF16 GGUF used for quantization.
- Runtime & Engine:
- charlie12345/ROCmFPX — High-performance ROCm/HIP llama.cpp fork with native FP4/WMMA quantization kernels tuned for AMD RDNA 3.5 architectures.
- nanashi66/ROCmFPX — Port & contribution (PR #113) backporting
spark2_5hybrid attention graph construction and pre-tokenizer fromggml-org/llama.cpp#27868.
- Calibration Dataset:
- eaddario/imatrix-calibration — High-quality calibration corpus by eaddario used for
llama-imatrixmulti-domain importance sampling.
- eaddario/imatrix-calibration — High-quality calibration corpus by eaddario used for
- Evaluation Framework:
- SeraphimSerapis/tool-eval-bench — Comprehensive 69-scenario agentic function-calling evaluation framework.
- Silicon & Compute Platform:
- AMD ROCm / HIP Stack — Native AMD ROCm 7.2 platform enabling accelerated WMMA Wave32 execution on
gfx1150andgfx1151.
- AMD ROCm / HIP Stack — Native AMD ROCm 7.2 platform enabling accelerated WMMA Wave32 execution on
- Downloads last month
- 240
4-bit