ZTFlynn commited on
Commit
5db3a88
Β·
verified Β·
1 Parent(s): b77c6fa

Add README.md

Browse files
Files changed (1) hide show
  1. README.md +57 -22
README.md CHANGED
@@ -23,7 +23,7 @@ language:
23
 
24
  # ZTFlynn/LFM2-1.2B-Extract-Cascadia-ternary3
25
 
26
- [`LiquidAI/LFM2-1.2B-Extract`](https://huggingface.co/LiquidAI/LFM2-1.2B-Extract) compressed to **747 MB**
27
  with Cascadia β€” a spline manifold plus per-band lookup tables at **0.60 bytes
28
  per weight** β€” and executable on CPU by a C runtime whose entire dependency
29
  list is libc, libm and libgomp.
@@ -32,8 +32,8 @@ list is libc, libm and libgomp.
32
  |---|---|
33
  | Base model | [`LiquidAI/LFM2-1.2B-Extract`](https://huggingface.co/LiquidAI/LFM2-1.2B-Extract) |
34
  | Parameters | 16 layers, hidden 2048 |
35
- | Checkpoint β†’ package | 2.23 GB β†’ **747 MB** (3.14x) |
36
- | Bits per weight | 5.09 |
37
  | Tensors compressed | 16 |
38
  | Architecture | 16 blocks, GQA 32q/8kv, gated short convolutions |
39
 
@@ -41,30 +41,65 @@ list is libc, libm and libgomp.
41
 
42
  | | perplexity |
43
  |---|---:|
44
- | `LiquidAI/LFM2-1.2B-Extract` (bf16) | 165.62 |
45
- | This package (ternary-3) | 175.84 |
46
- | **Result** | **+6.17% perplexity** (95% CI [1.0468x, 1.0768x], t = +8.31) |
47
 
48
- 8,176 paired tokens, FineWeb-Edu, 512-token windows. Scoring both
49
- models on identical tokens and comparing per token cuts the standard error
50
- 6.0x versus two independent means, which is what makes this
51
- resolution achievable.
52
 
53
- A measured cost. At this size the ternary-3 rate is a real trade: 3.14x compression for 6.2% higher perplexity. Larger models in this family pay far less β€” see the table below.
 
 
 
54
 
55
- ### How compression cost scales with model size
56
 
57
- Measured across the family with the same corpus and method:
58
 
59
- | model | parameters | perplexity cost |
60
- |---|---:|---|
61
- | LFM2.5-230M | 0.23B | +7.7% |
62
- | LFM2-350M | 0.35B | +3.5% |
63
- | LFM2-24B-A2B | 24B | no detectable cost (< 0.3%) |
64
 
65
- Redundancy grows faster than the format's error, so larger models compress
66
- more nearly losslessly. Below roughly 350M parameters the ternary-3 rate
67
- becomes a visible trade rather than a free one.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
68
 
69
  ## Usage
70
 
@@ -109,7 +144,7 @@ Greedy, generated to natural completion.
109
 
110
  | file | size |
111
  |---|---|
112
- | `weights.bin` | 745 MB |
113
  | `manifest.json` | per-tensor geometry and offsets |
114
  | `aux.bin` | RMSNorm scales, conv kernels, architecture constants |
115
  | `tokenizer.bin` | vocabulary, merges, Unicode tables |
 
23
 
24
  # ZTFlynn/LFM2-1.2B-Extract-Cascadia-ternary3
25
 
26
+ [`LiquidAI/LFM2-1.2B-Extract`](https://huggingface.co/LiquidAI/LFM2-1.2B-Extract) compressed to **774 MB**
27
  with Cascadia β€” a spline manifold plus per-band lookup tables at **0.60 bytes
28
  per weight** β€” and executable on CPU by a C runtime whose entire dependency
29
  list is libc, libm and libgomp.
 
32
  |---|---|
33
  | Base model | [`LiquidAI/LFM2-1.2B-Extract`](https://huggingface.co/LiquidAI/LFM2-1.2B-Extract) |
34
  | Parameters | 16 layers, hidden 2048 |
35
+ | Checkpoint β†’ package | 2.23 GB β†’ **774 MB** (3.03x) |
36
+ | Bits per weight | 5.28 |
37
  | Tensors compressed | 16 |
38
  | Architecture | 16 blocks, GQA 32q/8kv, gated short convolutions |
39
 
 
41
 
42
  | | perplexity |
43
  |---|---:|
44
+ | `LiquidAI/LFM2-1.2B-Extract` (bf16) | 265.77 |
45
+ | This package (ternary-3) | 267.71 |
46
+ | **Result** | **no detectable cost** (95% CI [0.9861x, 1.0329x], t = +0.78) |
47
 
48
+ 16,352 paired tokens of FineWeb-Edu in **31 independent
49
+ 512-token windows**. Both models score identical tokens and are
50
+ compared per token, which cuts the standard error 17.0x versus
51
+ two independent means.
52
 
53
+ The window is the unit of inference, not the token: tokens inside one window
54
+ share a context, and counting them as independent samples inflates the
55
+ t-statistic several-fold. Resolving a difference of a few percent takes
56
+ hundreds of windows.
57
 
58
+ The difference does not resolve. Stated precisely: this measurement bounds the perplexity change to within **3.3%** and cannot distinguish it from zero β€” which is a limit on the evidence, not a proof that the cost is zero. Reconstruction fidelity below is measured directly and carries no such uncertainty.
59
 
60
+ ### Reconstruction fidelity
61
 
62
+ Perplexity measures how good a model is on a corpus, not how faithful a copy
63
+ is, and the two disagree here: models compressed to identical reconstruction
64
+ error differ by 16 percentage points of measured perplexity. Fidelity has no
65
+ sampling uncertainty and no dependence on corpus domain, so it is measured
66
+ directly and reported alongside.
67
 
68
+ | | |
69
+ |---|---:|
70
+ | Relative L2 error vs the bf16 checkpoint | **0.0553** |
71
+ | Systematic gain (1.0000 is faithful) | 0.9993 |
72
+ | Measured over | 93 of 93 tensors, 100% of parameters |
73
+
74
+ By tensor class:
75
+
76
+ | class | rel L2 | share of model |
77
+ |---|---:|---:|
78
+ | linear | 0.0576 | 1,036M params |
79
+ | embedding | 0.0265 | 134M params |
80
+
81
+ The tied embedding, the tensor whose error reaches the logits undamped, reconstructs at **0.0265**.
82
+
83
+ ### Where the compression cost comes from
84
+
85
+ The cost is concentrated in one tensor. The tied embedding, which also
86
+ serves as `lm_head`, is compressed by a single global codebook β€” no bands,
87
+ no spline manifold, no exact outliers β€” while every linear tensor gets 32
88
+ bands, a spline, and 0.5% of its weights kept exact. Its error is the only
89
+ error in the model that reaches the logits with nothing downstream to
90
+ absorb it.
91
+
92
+ Measured on LFM2-350M, relative L2 reconstruction error:
93
+
94
+ | tensor | codebook | rel L2 |
95
+ |---|---:|---:|
96
+ | tied embedding, 27 entries | 27 | 0.078 |
97
+ | tied embedding, 81 entries | 81 | **0.027** |
98
+ | a typical linear (32 bands x 27) | 864 | 0.057 |
99
+
100
+ At 27 entries the embedding is the worst-reconstructed tensor in the model.
101
+ This package uses **81** entries for it, which costs about 6% in size and
102
+ makes it the best-reconstructed tensor instead.
103
 
104
  ## Usage
105
 
 
144
 
145
  | file | size |
146
  |---|---|
147
+ | `weights.bin` | 772 MB |
148
  | `manifest.json` | per-tensor geometry and offsets |
149
  | `aux.bin` | RMSNorm scales, conv kernels, architecture constants |
150
  | `tokenizer.bin` | vocabulary, merges, Unicode tables |