ZTFlynn commited on
Commit
729f0ce
·
verified ·
1 Parent(s): 53aabf9

Refresh card with final measurements

Browse files
Files changed (1) hide show
  1. README.md +24 -0
README.md CHANGED
@@ -45,6 +45,30 @@ list is libc, libm and libgomp.
45
  | This package (ternary-3) | 248.74 |
46
  | **Result** | **+1.10% perplexity** — marginal (95% CI [1.0003x, 1.0219x], t = +2.01) |
47
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
48
  117,530 paired tokens of FineWeb-Edu in **229 independent
49
  512-token windows**. Both models score identical tokens and are
50
  compared per token, which cuts the standard error 10.3x versus
 
45
  | This package (ternary-3) | 248.74 |
46
  | **Result** | **+1.10% perplexity** — marginal (95% CI [1.0003x, 1.0219x], t = +2.01) |
47
 
48
+
49
+ ### Against the 4-bit formats the hardware runs natively
50
+
51
+ NVIDIA Blackwell executes NVFP4 and MXFP4 in the tensor cores with no software
52
+ decode. Cascadia derives a per-element band instead of storing a per-block
53
+ scale, which costs real decode time -- so the question is whether it buys
54
+ anything. Every format below was scored through one evaluation path on the
55
+ same 119,574 tokens, with the same window-clustered pairing.
56
+
57
+ | format | bytes/weight | perplexity | vs this package |
58
+ |---|---:|---:|---|
59
+ | bf16 (uncompressed) | 2.0000 | 242.65 | — |
60
+ | **Cascadia ternary-3** | 0.6746 | 244.80 | — |
61
+ | NVFP4 (e2m1 + e4m3 / 16) | 0.5625 | 266.74 | **1.090x worse** (t = +7.80) |
62
+ | int4 + fp16 scale / 32 | 0.5625 | 270.84 | **1.106x worse** (t = +9.55) |
63
+ | MXFP4 (e2m1 + ue8m0 / 32) | 0.5312 | 314.84 | **1.286x worse** (t = +22.44) |
64
+
65
+ Cascadia is 20% larger than NVFP4 and resolvably better than all three, while its own gap to bf16 is *not* statistically resolvable on this corpus (t = +1.52). The result is size-dependent and does not generalise: on the smaller LFM2.5-230M, int4 with a per-32 fp16 scale beats Cascadia while being smaller.
66
+
67
+ Reproduce with `tools/format_shootout.py`. Its `cascadia` arm reconstructs
68
+ this package and is checked against the C runtime's own perplexity before any
69
+ number is reported; a mismatch aborts rather than prints.
70
+
71
+
72
  117,530 paired tokens of FineWeb-Edu in **229 independent
73
  512-token windows**. Both models score identical tokens and are
74
  compared per token, which cuts the standard error 10.3x versus