File size: 4,647 Bytes
fbfb183
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
---
license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-8B-A1B/blob/main/LICENSE
base_model: LiquidAI/LFM2.5-8B-A1B
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
  - gguf
  - llama.cpp
  - rocm
  - amd
  - rocmfp4
  - rocmfpx
  - strix-halo
  - amd-strix-halo
  - gfx1151
  - ryzen-ai-max
  - ryzen-ai-max-395
  - radeon-8060s
  - lfm2
  - liquid-ai
  - quantized
---

# LFM2.5-8B-A1B (LEAN) β€” ROCmFP4 for AMD Strix Halo (gfx1151)

> βœ… **the first ROCmFP4 build of any LFM2.5 checkpoint**
>
> *Checked 2026-08-22 against every public GGUF of this model. All existing builds
> (LiquidAI's own, unsloth, and others) ship standard k-quants. ROCmFP4 is a runtime tensor
> format that exists only in the [ROCmFPX](https://github.com/charlie12345/ROCmFPX) fork of
> llama.cpp. Repository-content comparison only β€” no third-party build was run or benchmarked here.*

A 4-bit ROCmFP4 quantisation of **LiquidAI/LFM2.5-8B-A1B** for AMD Ryzen AI Max+ 395 / Radeon 8060S / gfx1151.

## The file

| | |
|---|---|
| ftype | `101` β€” `Q4_0_ROCMFP4_LEAN` |
| size | **4,809,862,624 bytes** (4.48 GiB) |
| architecture | `lfm2moe` |
| tensors | 256 |
| context | 128,000 |
| token embedding | `Q5_K` |

Type histogram, read from the finished file:

```
ROCmFP4 x132, F32 x123, Q5_K x1
```

The LEAN (101) and COHERENT (102) tiers differ only in the token-embedding type β€”
`Q5_K` for LEAN, `Q6_K` for COHERENT. All other tensors are identical. This model ties its output
projection to `token_embd.weight`, so there is no separate `output.weight` to protect.

## Measured throughput

AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), ROCm 7.13.0, 125 GB unified memory, idle box.
`llama-cli -ngl 999 -fa on -c 512 -n 64 --temp 0 --seed 1234`:

| | generation |
|---|---:|
| this file | **147.1 t/s** |

A separate 3-repetition benchmark at `-c 2048 -n 512` measured **137.3 t/s** for this
checkpoint with no drafter.

## ⚠️ DSpark speculative decoding is a NET LOSS on this hardware β€” do not use it

LiquidAI publishes a DSpark speculator for this model. **We measured it and it makes generation
slower**, so no ROCmFP4 draft is published here.

| config | generation | effect |
|---|---:|---:|
| no drafter | 137.3 t/s | β€” |
| `--spec-type draft-dspark --spec-draft-n-max 8` | 85.1 t/s | **-38.0%** |

Mean accepted length was **2.71** (block size 9). Across all three LFM2.5 sizes the result was
consistently negative: βˆ’28.4% (1.2B), βˆ’19.1% (2.6B), βˆ’38.0% (8B-A1B).

Two causes were identified, both in the runtime rather than the weights:

1. `lfm2.cpp` / `lfm2moe.cpp` do not populate `t_layer_inp[]`, so `draft-dspark` aborts on
   `GGML_ASSERT(t_layer_inp[il] != nullptr)` out of the box. A one-line patch
   (`res->t_layer_inp[il] = prev_cur;`) makes it run.
2. With that fixed, llama.cpp reports *recurrent state rollback is not compatible with
   'draft-dspark'* and falls back to a checkpoint path that is **not bit-exact** for LFM2's
   recurrent state β€” DSpark output diverges from greedy target output (reproducible 3/3).

An off-by-one in the target-layer mapping was ruled out: forcing
`LLAMA_DFLASH_TARGET_LAYER_OFFSET=-1` produced a *worse* accepted length (2.22), confirming the
converter's `+1` convention is correct.

**DSpark on LFM2.5 needs real recurrent-state rollback support before any draft is worth shipping.**

## Requirements

This file uses the ROCmFP4 tensor format, which exists only in the
[ROCmFPX](https://github.com/charlie12345/ROCmFPX) fork of llama.cpp. Stock llama.cpp will not
load it.

```bash
llama-cli -m LFM2.5-8B-A1B-Q4_0_ROCMFP4_LEAN.gguf \
  -ngl 999 -fa on -c 2048 -n 512 \
  -p "The history of mathematics begins in ancient times. One of the earliest known"
```

## Sample output

Continuation from `"The history of mathematics begins in ancient times. One of the earliest known"`:

> [Start thinking]
> The user gave a partial sentence: "The history of mathematics begins in ancient times. One of the earliest known ...". They likely want continuation.

## Not measured

Perplexity is not published for this build; quality evidence here is the coherence check above and
the tensor-level audit. Long-context behaviour at the full 128,000-token window was not tested.

## Provenance

Converted from `LiquidAI/LFM2.5-8B-A1B` at revision `b9aebfcbe28b6cb374042f495d733037550ab146` to F16 GGUF using upstream
[llama.cpp](https://github.com/ggml-org/llama.cpp) at `e85caa81ea2b65797396018c179b87ad61fa38ab`, then quantised to ftype 101
with the ROCmFPX fork (`feature/dspark-v2`). Licence inherited from the base model.