--- license: other license_name: lfm1.0 license_link: https://huggingface.co/LiquidAI/LFM2.5-8B-A1B/blob/main/LICENSE base_model: LiquidAI/LFM2.5-8B-A1B base_model_relation: quantized pipeline_tag: text-generation library_name: gguf tags: - gguf - llama.cpp - rocm - amd - rocmfp4 - rocmfpx - strix-halo - amd-strix-halo - gfx1151 - ryzen-ai-max - ryzen-ai-max-395 - radeon-8060s - lfm2 - liquid-ai - quantized --- # LFM2.5-8B-A1B (LEAN) — ROCmFP4 for AMD Strix Halo (gfx1151) > ✅ **the first ROCmFP4 build of any LFM2.5 checkpoint** > > *Checked 2026-08-22 against every public GGUF of this model. All existing builds > (LiquidAI's own, unsloth, and others) ship standard k-quants. ROCmFP4 is a runtime tensor > format that exists only in the [ROCmFPX](https://github.com/charlie12345/ROCmFPX) fork of > llama.cpp. Repository-content comparison only — no third-party build was run or benchmarked here.* A 4-bit ROCmFP4 quantisation of **LiquidAI/LFM2.5-8B-A1B** for AMD Ryzen AI Max+ 395 / Radeon 8060S / gfx1151. ## The file | | | |---|---| | ftype | `101` — `Q4_0_ROCMFP4_LEAN` | | size | **4,809,862,624 bytes** (4.48 GiB) | | architecture | `lfm2moe` | | tensors | 256 | | context | 128,000 | | token embedding | `Q5_K` | Type histogram, read from the finished file: ``` ROCmFP4 x132, F32 x123, Q5_K x1 ``` The LEAN (101) and COHERENT (102) tiers differ only in the token-embedding type — `Q5_K` for LEAN, `Q6_K` for COHERENT. All other tensors are identical. This model ties its output projection to `token_embd.weight`, so there is no separate `output.weight` to protect. ## Measured throughput AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), ROCm 7.13.0, 125 GB unified memory, idle box. `llama-cli -ngl 999 -fa on -c 512 -n 64 --temp 0 --seed 1234`: | | generation | |---|---:| | this file | **147.1 t/s** | A separate 3-repetition benchmark at `-c 2048 -n 512` measured **137.3 t/s** for this checkpoint with no drafter. ## ⚠️ DSpark speculative decoding is a NET LOSS on this hardware — do not use it LiquidAI publishes a DSpark speculator for this model. **We measured it and it makes generation slower**, so no ROCmFP4 draft is published here. | config | generation | effect | |---|---:|---:| | no drafter | 137.3 t/s | — | | `--spec-type draft-dspark --spec-draft-n-max 8` | 85.1 t/s | **-38.0%** | Mean accepted length was **2.71** (block size 9). Across all three LFM2.5 sizes the result was consistently negative: −28.4% (1.2B), −19.1% (2.6B), −38.0% (8B-A1B). Two causes were identified, both in the runtime rather than the weights: 1. `lfm2.cpp` / `lfm2moe.cpp` do not populate `t_layer_inp[]`, so `draft-dspark` aborts on `GGML_ASSERT(t_layer_inp[il] != nullptr)` out of the box. A one-line patch (`res->t_layer_inp[il] = prev_cur;`) makes it run. 2. With that fixed, llama.cpp reports *recurrent state rollback is not compatible with 'draft-dspark'* and falls back to a checkpoint path that is **not bit-exact** for LFM2's recurrent state — DSpark output diverges from greedy target output (reproducible 3/3). An off-by-one in the target-layer mapping was ruled out: forcing `LLAMA_DFLASH_TARGET_LAYER_OFFSET=-1` produced a *worse* accepted length (2.22), confirming the converter's `+1` convention is correct. **DSpark on LFM2.5 needs real recurrent-state rollback support before any draft is worth shipping.** ## Requirements This file uses the ROCmFP4 tensor format, which exists only in the [ROCmFPX](https://github.com/charlie12345/ROCmFPX) fork of llama.cpp. Stock llama.cpp will not load it. ```bash llama-cli -m LFM2.5-8B-A1B-Q4_0_ROCMFP4_LEAN.gguf \ -ngl 999 -fa on -c 2048 -n 512 \ -p "The history of mathematics begins in ancient times. One of the earliest known" ``` ## Sample output Continuation from `"The history of mathematics begins in ancient times. One of the earliest known"`: > [Start thinking] > The user gave a partial sentence: "The history of mathematics begins in ancient times. One of the earliest known ...". They likely want continuation. ## Not measured Perplexity is not published for this build; quality evidence here is the coherence check above and the tensor-level audit. Long-context behaviour at the full 128,000-token window was not tested. ## Provenance Converted from `LiquidAI/LFM2.5-8B-A1B` at revision `b9aebfcbe28b6cb374042f495d733037550ab146` to F16 GGUF using upstream [llama.cpp](https://github.com/ggml-org/llama.cpp) at `e85caa81ea2b65797396018c179b87ad61fa38ab`, then quantised to ftype 101 with the ROCmFPX fork (`feature/dspark-v2`). Licence inherited from the base model.