Text Generation
GGUF
llama.cpp
rocm
amd
rocmfp4
rocmfpx
strix-halo
amd-strix-halo
gfx1151
ryzen-ai-max
ryzen-ai-max-395
radeon-8060s
lfm2
liquid-ai
quantized
conversational
Instructions to use kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: llama cli -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: llama cli -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: ./llama-cli -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: ./build/bin/llama-cli -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
Use Docker
docker model run hf.co/kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
- LM Studio
- Jan
- vLLM
How to use kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
- Ollama
How to use kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF with Ollama:
ollama run hf.co/kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
- Unsloth Desktop
- Pi
How to use kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF with Docker Model Runner:
docker model run hf.co/kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
- Lemonade
How to use kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
Run and chat with the model
lemonade run user.LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF-Q4_0_ROCMFP
List all available models
lemonade list
- Hermes Agent
How to use kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "kingjones777/LFM2.5-8B-A1B-ROCmFP4-LEAN-GGUF:Q4_0_ROCMFP" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 4,647 Bytes
fbfb183 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 | ---
license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-8B-A1B/blob/main/LICENSE
base_model: LiquidAI/LFM2.5-8B-A1B
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- llama.cpp
- rocm
- amd
- rocmfp4
- rocmfpx
- strix-halo
- amd-strix-halo
- gfx1151
- ryzen-ai-max
- ryzen-ai-max-395
- radeon-8060s
- lfm2
- liquid-ai
- quantized
---
# LFM2.5-8B-A1B (LEAN) β ROCmFP4 for AMD Strix Halo (gfx1151)
> β
**the first ROCmFP4 build of any LFM2.5 checkpoint**
>
> *Checked 2026-08-22 against every public GGUF of this model. All existing builds
> (LiquidAI's own, unsloth, and others) ship standard k-quants. ROCmFP4 is a runtime tensor
> format that exists only in the [ROCmFPX](https://github.com/charlie12345/ROCmFPX) fork of
> llama.cpp. Repository-content comparison only β no third-party build was run or benchmarked here.*
A 4-bit ROCmFP4 quantisation of **LiquidAI/LFM2.5-8B-A1B** for AMD Ryzen AI Max+ 395 / Radeon 8060S / gfx1151.
## The file
| | |
|---|---|
| ftype | `101` β `Q4_0_ROCMFP4_LEAN` |
| size | **4,809,862,624 bytes** (4.48 GiB) |
| architecture | `lfm2moe` |
| tensors | 256 |
| context | 128,000 |
| token embedding | `Q5_K` |
Type histogram, read from the finished file:
```
ROCmFP4 x132, F32 x123, Q5_K x1
```
The LEAN (101) and COHERENT (102) tiers differ only in the token-embedding type β
`Q5_K` for LEAN, `Q6_K` for COHERENT. All other tensors are identical. This model ties its output
projection to `token_embd.weight`, so there is no separate `output.weight` to protect.
## Measured throughput
AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), ROCm 7.13.0, 125 GB unified memory, idle box.
`llama-cli -ngl 999 -fa on -c 512 -n 64 --temp 0 --seed 1234`:
| | generation |
|---|---:|
| this file | **147.1 t/s** |
A separate 3-repetition benchmark at `-c 2048 -n 512` measured **137.3 t/s** for this
checkpoint with no drafter.
## β οΈ DSpark speculative decoding is a NET LOSS on this hardware β do not use it
LiquidAI publishes a DSpark speculator for this model. **We measured it and it makes generation
slower**, so no ROCmFP4 draft is published here.
| config | generation | effect |
|---|---:|---:|
| no drafter | 137.3 t/s | β |
| `--spec-type draft-dspark --spec-draft-n-max 8` | 85.1 t/s | **-38.0%** |
Mean accepted length was **2.71** (block size 9). Across all three LFM2.5 sizes the result was
consistently negative: β28.4% (1.2B), β19.1% (2.6B), β38.0% (8B-A1B).
Two causes were identified, both in the runtime rather than the weights:
1. `lfm2.cpp` / `lfm2moe.cpp` do not populate `t_layer_inp[]`, so `draft-dspark` aborts on
`GGML_ASSERT(t_layer_inp[il] != nullptr)` out of the box. A one-line patch
(`res->t_layer_inp[il] = prev_cur;`) makes it run.
2. With that fixed, llama.cpp reports *recurrent state rollback is not compatible with
'draft-dspark'* and falls back to a checkpoint path that is **not bit-exact** for LFM2's
recurrent state β DSpark output diverges from greedy target output (reproducible 3/3).
An off-by-one in the target-layer mapping was ruled out: forcing
`LLAMA_DFLASH_TARGET_LAYER_OFFSET=-1` produced a *worse* accepted length (2.22), confirming the
converter's `+1` convention is correct.
**DSpark on LFM2.5 needs real recurrent-state rollback support before any draft is worth shipping.**
## Requirements
This file uses the ROCmFP4 tensor format, which exists only in the
[ROCmFPX](https://github.com/charlie12345/ROCmFPX) fork of llama.cpp. Stock llama.cpp will not
load it.
```bash
llama-cli -m LFM2.5-8B-A1B-Q4_0_ROCMFP4_LEAN.gguf \
-ngl 999 -fa on -c 2048 -n 512 \
-p "The history of mathematics begins in ancient times. One of the earliest known"
```
## Sample output
Continuation from `"The history of mathematics begins in ancient times. One of the earliest known"`:
> [Start thinking]
> The user gave a partial sentence: "The history of mathematics begins in ancient times. One of the earliest known ...". They likely want continuation.
## Not measured
Perplexity is not published for this build; quality evidence here is the coherence check above and
the tensor-level audit. Long-context behaviour at the full 128,000-token window was not tested.
## Provenance
Converted from `LiquidAI/LFM2.5-8B-A1B` at revision `b9aebfcbe28b6cb374042f495d733037550ab146` to F16 GGUF using upstream
[llama.cpp](https://github.com/ggml-org/llama.cpp) at `e85caa81ea2b65797396018c179b87ad61fa38ab`, then quantised to ftype 101
with the ROCmFPX fork (`feature/dspark-v2`). Licence inherited from the base model.
|