Refresh card with final measurements
Browse files
README.md
CHANGED
|
@@ -35,6 +35,7 @@ list is libc, libm and libgomp.
|
|
| 35 |
| Checkpoint → package | 2.23 GB → **774 MB** (3.03x) |
|
| 36 |
| Bits per weight | 5.28 |
|
| 37 |
| Tensors compressed | 16 |
|
|
|
|
| 38 |
| Architecture | 16 blocks, GQA 32q/8kv, gated short convolutions |
|
| 39 |
|
| 40 |
## Quality
|
|
@@ -46,6 +47,21 @@ list is libc, libm and libgomp.
|
|
| 46 |
| **Result** | **no detectable cost** (95% CI [0.9899x, 1.0104x], t = +0.02) |
|
| 47 |
|
| 48 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 49 |
65,408 paired tokens of FineWeb-Edu in **127 independent
|
| 50 |
512-token windows**. Both models score identical tokens and are
|
| 51 |
compared per token, which cuts the standard error 15.7x versus
|
|
|
|
| 35 |
| Checkpoint → package | 2.23 GB → **774 MB** (3.03x) |
|
| 36 |
| Bits per weight | 5.28 |
|
| 37 |
| Tensors compressed | 16 |
|
| 38 |
+
| Resident memory | **870 MB** on CPU, 919 MB on GPU — 1.12x the package |
|
| 39 |
| Architecture | 16 blocks, GQA 32q/8kv, gated short convolutions |
|
| 40 |
|
| 41 |
## Quality
|
|
|
|
| 47 |
| **Result** | **no detectable cost** (95% CI [0.9899x, 1.0104x], t = +0.02) |
|
| 48 |
|
| 49 |
|
| 50 |
+
## Running it
|
| 51 |
+
|
| 52 |
+
Measured on a Jetson AGX Thor: CPU is 12 threads of a 14-core Arm part, GPU is
|
| 53 |
+
the integrated Blackwell. Decode is the median of five runs; prefill is from a
|
| 54 |
+
162-token prompt. Both backends produce **byte-identical** output.
|
| 55 |
+
|
| 56 |
+
| | CPU (12 threads) | GPU |
|
| 57 |
+
|---|---|---|
|
| 58 |
+
| decode | 2.43 tok/s | 24.66 |
|
| 59 |
+
| prefill | 16.7 tok/s | 90.0 |
|
| 60 |
+
|
| 61 |
+
Load takes 1.34 s on CPU and 0.82 s on GPU,
|
| 62 |
+
the latter including the host-to-device upload.
|
| 63 |
+
|
| 64 |
+
|
| 65 |
65,408 paired tokens of FineWeb-Edu in **127 independent
|
| 66 |
512-token windows**. Both models score identical tokens and are
|
| 67 |
compared per token, which cuts the standard error 15.7x versus
|