ZTFlynn/LFM2.5-8B-A1B-Cascadia-ternary3

LiquidAI/LFM2.5-8B-A1B compressed to 5.49 GB with Cascadia β€” a spline manifold plus per-band lookup tables at 0.647 bytes per weight β€” and executable on CPU by a C runtime whose entire dependency list is libc, libm and libgomp.

Base model LiquidAI/LFM2.5-8B-A1B
Parameters 24 layers, hidden 2048
Checkpoint β†’ package 16.15 GB β†’ 5.49 GB (3.09x)
Bits per weight 5.18
Tensors compressed 24
Resident memory 6083 MB on CPU β€” 1.08x the package
Architecture 24 blocks, GQA 32q/8kv, gated short convolutions, 32 experts top-4 above 2 dense layers

Quality

perplexity
LiquidAI/LFM2.5-8B-A1B (bf16) 63.39
This package (ternary-3) 65.03
Result +2.58% perplexity β€” marginal (95% CI [1.0045x, 1.0455x], t = +2.40)

Running it

Measured on a Jetson AGX Thor: CPU is 12 threads of a 14-core Arm part, GPU is the integrated Blackwell. Decode is the median of five runs; prefill is from a 162-token prompt. Both backends produce byte-identical output.

CPU (12 threads) GPU
decode 1.72 tok/s n/a
prefill 10.0 tok/s n/a

Load takes 10.05 s on CPU and None s on GPU, the latter including the host-to-device upload.

32,704 paired tokens of FineWeb-Edu in 63 independent 512-token windows. Both models score identical tokens and are compared per token, which cuts the standard error 6.1x versus two independent means.

The window is the unit of inference, not the token: tokens inside one window share a context, and counting them as independent samples inflates the t-statistic several-fold. Resolving a difference of a few percent takes hundreds of windows.

A cost near the edge of what this sample resolves: t = +2.40 clears the conventional 1.96 on its own, but 13 models were measured together and holding family-wise error at 5% across them needs |t| > 2.89. Read 2.6% as the point estimate, with the interval above carrying the real uncertainty.

The corpus is general web text (FineWeb-Edu). How much a given model is affected depends on its own sensitivity and on how close that text is to its domain, so this figure is specific to both. Two models in this family compressed to the same 0.055 reconstruction error measured +13.4% and βˆ’2.4% here. The fidelity figure below describes the compression itself and does not vary that way.

Reconstruction fidelity

Perplexity measures how good a model is on a corpus, not how faithful a copy is, and the two disagree here: models compressed to identical reconstruction error differ by 16 percentage points of measured perplexity. Fidelity has no sampling uncertainty and no dependence on corpus domain, so it is measured directly and reported alongside.

Relative L2 error vs the bf16 checkpoint 0.0548
Systematic gain (1.0000 is faithful) 0.9994
Measured over 256 of 1497 tensors, 20% of parameters

By tensor class:

class rel L2 share of model
expert 0.0570 7,751M params
linear 0.0618 454M params
embedding 0.0224 262M params

Experts are homogeneous, so a stride across them estimates the population; the figure above extrapolates each class from its sample to that class's full parameter count.

The tied embedding, the tensor whose error reaches the logits undamped, reconstructs at 0.0224.

Where the compression cost comes from

The cost is concentrated in one tensor. The tied embedding, which also serves as lm_head, is compressed by a single global codebook β€” no bands, no spline manifold, no exact outliers β€” while every linear tensor gets 32 bands, a spline, and 0.5% of its weights kept exact. Its error is the only error in the model that reaches the logits with nothing downstream to absorb it.

Measured on LFM2-350M, relative L2 reconstruction error:

tensor codebook rel L2
tied embedding, 27 entries 27 0.078
tied embedding, 81 entries 81 0.027
a typical linear (32 bands x 27) 864 0.057

At 27 entries the embedding is the worst-reconstructed tensor in the model. This package uses 81 entries for it, which costs about 6% in size and makes it the best-reconstructed tensor instead.

Usage

Executed by the Cascadia C runtime. This is a compressed package, not a transformers checkpoint.

git clone https://github.com/EntroMorphic/cassie && cd cassie
cmake -S src/c -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build -j

huggingface-cli download ZTFlynn/LFM2.5-8B-A1B-Cascadia-ternary3 --local-dir ./pkg
./build/cascadia_generate ./pkg 512 --chat "Explain gradient descent."

Sampling is --temp / --top-k / --top-p / --seed; the default is greedy and seed-reproducible. Generation stops at <|im_end|>, so max_new is a ceiling.

This package is Mixture-of-Experts, which the CUDA backend does not execute yet; it runs on the C runtime above.

Python

from transformers import AutoModelForCausalLM
from cascadia import load_compressed

model = AutoModelForCausalLM.from_pretrained("LiquidAI/LFM2.5-8B-A1B", dtype="bfloat16")
model, stats = load_compressed(model, "./pkg", model_id="LiquidAI/LFM2.5-8B-A1B")

Sample output

Prompt: "How many eggs are in a baker's dozen?"


Greedy, generated to natural completion.

Package contents

file size
weights.bin 5.49 GB
manifest.json per-tensor geometry and offsets
aux.bin RMSNorm scales, conv kernels, architecture constants
tokenizer.bin vocabulary, merges, Unicode tables

Format specified in docs/package_format.md and machine-verified against every package.

How it works

A B-spline surface is fitted to each weight matrix to capture large-scale structure. Each weight is assigned to one of 32 bands by its spline value, and a k-means codebook is learned per band over the residuals. The top 0.5% of errors are kept exactly as f32. Codebook indices pack in base 3, five trits per byte, since 3⁡ = 243 fits a byte.

Reconstruction is W = spline(j,c) + codebook[band][index], evaluated inside the matvec so no dense weight matrix is ever built. Because the spline carries dynamic range, the residual tables need no per-block scale factors.

Limitations

  • Runs under the Cascadia C runtime rather than transformers directly.
  • The runtime executes ternary-3 packages; other presets convert but are not yet supported by the kernel.
  • Batch-1 CPU inference, suited to edge and batch workloads.
  • Greedy and sampled decoding; no beam search.

Acknowledgements

Deeply inspired by Magneato/deepseek-r1-qwen-7b-lutc, which demonstrated LUT-cascade compression of a 7B model at 5.45 bits per weight. The Guanaco LUT cascade β€” no-scale residuals, variable bit rate, and f32 outlier preservation β€” is the foundation this builds on. Cascadia adds a spline manifold for band selection and a Harmonic Collapse step that removes per-block scale factors entirely. Our thanks to Magneato for publishing both the approach and the weights that made it concrete.

Base model by Liquid AI, used under the LFM Open License.

Citation

@software{cascadia,
  title  = {Cascadia: Spline Manifold LUT Compression for Language Models},
  author = {Josserand-Austin, Tripp},
  year   = {2026},
  url    = {https://github.com/EntroMorphic/cassie}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ZTFlynn/LFM2.5-8B-A1B-Cascadia-ternary3

Quantized
(95)
this model