ZTFlynn commited on
Commit
f3429cb
·
verified ·
1 Parent(s): 82237e0

Refresh card with final measurements

Browse files
Files changed (1) hide show
  1. README.md +16 -0
README.md CHANGED
@@ -35,6 +35,7 @@ list is libc, libm and libgomp.
35
  | Checkpoint → package | 2.23 GB → **774 MB** (3.03x) |
36
  | Bits per weight | 5.28 |
37
  | Tensors compressed | 16 |
 
38
  | Architecture | 16 blocks, GQA 32q/8kv, gated short convolutions |
39
 
40
  ## Quality
@@ -46,6 +47,21 @@ list is libc, libm and libgomp.
46
  | **Result** | **+2.04% perplexity** — marginal (95% CI [1.0040x, 1.0320x], t = +2.53) |
47
 
48
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
49
  16,352 paired tokens of FineWeb-Edu in **31 independent
50
  512-token windows**. Both models score identical tokens and are
51
  compared per token, which cuts the standard error 16.3x versus
 
35
  | Checkpoint → package | 2.23 GB → **774 MB** (3.03x) |
36
  | Bits per weight | 5.28 |
37
  | Tensors compressed | 16 |
38
+ | Resident memory | **870 MB** on CPU, 919 MB on GPU — 1.12x the package |
39
  | Architecture | 16 blocks, GQA 32q/8kv, gated short convolutions |
40
 
41
  ## Quality
 
47
  | **Result** | **+2.04% perplexity** — marginal (95% CI [1.0040x, 1.0320x], t = +2.53) |
48
 
49
 
50
+ ## Running it
51
+
52
+ Measured on a Jetson AGX Thor: CPU is 12 threads of a 14-core Arm part, GPU is
53
+ the integrated Blackwell. Decode is the median of five runs; prefill is from a
54
+ 162-token prompt. Both backends produce **byte-identical** output.
55
+
56
+ | | CPU (12 threads) | GPU |
57
+ |---|---|---|
58
+ | decode | 2.43 tok/s | 24.69 |
59
+ | prefill | 14.8 tok/s | 89.9 |
60
+
61
+ Load takes 0.83 s on CPU and 0.82 s on GPU,
62
+ the latter including the host-to-device upload.
63
+
64
+
65
  16,352 paired tokens of FineWeb-Edu in **31 independent
66
  512-token windows**. Both models score identical tokens and are
67
  compared per token, which cuts the standard error 16.3x versus