Refresh card with final measurements
Browse files
README.md
CHANGED
|
@@ -45,6 +45,30 @@ list is libc, libm and libgomp.
|
|
| 45 |
| This package (ternary-3) | 248.74 |
|
| 46 |
| **Result** | **+1.10% perplexity** — marginal (95% CI [1.0003x, 1.0219x], t = +2.01) |
|
| 47 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 48 |
117,530 paired tokens of FineWeb-Edu in **229 independent
|
| 49 |
512-token windows**. Both models score identical tokens and are
|
| 50 |
compared per token, which cuts the standard error 10.3x versus
|
|
|
|
| 45 |
| This package (ternary-3) | 248.74 |
|
| 46 |
| **Result** | **+1.10% perplexity** — marginal (95% CI [1.0003x, 1.0219x], t = +2.01) |
|
| 47 |
|
| 48 |
+
|
| 49 |
+
### Against the 4-bit formats the hardware runs natively
|
| 50 |
+
|
| 51 |
+
NVIDIA Blackwell executes NVFP4 and MXFP4 in the tensor cores with no software
|
| 52 |
+
decode. Cascadia derives a per-element band instead of storing a per-block
|
| 53 |
+
scale, which costs real decode time -- so the question is whether it buys
|
| 54 |
+
anything. Every format below was scored through one evaluation path on the
|
| 55 |
+
same 119,574 tokens, with the same window-clustered pairing.
|
| 56 |
+
|
| 57 |
+
| format | bytes/weight | perplexity | vs this package |
|
| 58 |
+
|---|---:|---:|---|
|
| 59 |
+
| bf16 (uncompressed) | 2.0000 | 242.65 | — |
|
| 60 |
+
| **Cascadia ternary-3** | 0.6746 | 244.80 | — |
|
| 61 |
+
| NVFP4 (e2m1 + e4m3 / 16) | 0.5625 | 266.74 | **1.090x worse** (t = +7.80) |
|
| 62 |
+
| int4 + fp16 scale / 32 | 0.5625 | 270.84 | **1.106x worse** (t = +9.55) |
|
| 63 |
+
| MXFP4 (e2m1 + ue8m0 / 32) | 0.5312 | 314.84 | **1.286x worse** (t = +22.44) |
|
| 64 |
+
|
| 65 |
+
Cascadia is 20% larger than NVFP4 and resolvably better than all three, while its own gap to bf16 is *not* statistically resolvable on this corpus (t = +1.52). The result is size-dependent and does not generalise: on the smaller LFM2.5-230M, int4 with a per-32 fp16 scale beats Cascadia while being smaller.
|
| 66 |
+
|
| 67 |
+
Reproduce with `tools/format_shootout.py`. Its `cascadia` arm reconstructs
|
| 68 |
+
this package and is checked against the C runtime's own perplexity before any
|
| 69 |
+
number is reported; a mismatch aborts rather than prints.
|
| 70 |
+
|
| 71 |
+
|
| 72 |
117,530 paired tokens of FineWeb-Edu in **229 independent
|
| 73 |
512-token windows**. Both models score identical tokens and are
|
| 74 |
compared per token, which cuts the standard error 10.3x versus
|