Refresh card with final measurements
Browse files
README.md
CHANGED
|
@@ -35,6 +35,7 @@ list is libc, libm and libgomp.
|
|
| 35 |
| Checkpoint → package | 676 MB → **241 MB** (2.96x) |
|
| 36 |
| Bits per weight | 5.40 |
|
| 37 |
| Tensors compressed | 16 |
|
|
|
|
| 38 |
| Architecture | 16 blocks, GQA 16q/8kv, gated short convolutions |
|
| 39 |
|
| 40 |
## Quality
|
|
@@ -69,6 +70,21 @@ this package and is checked against the C runtime's own perplexity before any
|
|
| 69 |
number is reported; a mismatch aborts rather than prints.
|
| 70 |
|
| 71 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 72 |
117,530 paired tokens of FineWeb-Edu in **229 independent
|
| 73 |
512-token windows**. Both models score identical tokens and are
|
| 74 |
compared per token, which cuts the standard error 10.3x versus
|
|
|
|
| 35 |
| Checkpoint → package | 676 MB → **241 MB** (2.96x) |
|
| 36 |
| Bits per weight | 5.40 |
|
| 37 |
| Tensors compressed | 16 |
|
| 38 |
+
| Resident memory | **298 MB** on CPU, 355 MB on GPU — 1.24x the package |
|
| 39 |
| Architecture | 16 blocks, GQA 16q/8kv, gated short convolutions |
|
| 40 |
|
| 41 |
## Quality
|
|
|
|
| 70 |
number is reported; a mismatch aborts rather than prints.
|
| 71 |
|
| 72 |
|
| 73 |
+
## Running it
|
| 74 |
+
|
| 75 |
+
Measured on a Jetson AGX Thor: CPU is 12 threads of a 14-core Arm part, GPU is
|
| 76 |
+
the integrated Blackwell. Decode is the median of five runs; prefill is from a
|
| 77 |
+
162-token prompt. Both backends produce **byte-identical** output.
|
| 78 |
+
|
| 79 |
+
| | CPU (12 threads) | GPU |
|
| 80 |
+
|---|---|---|
|
| 81 |
+
| decode | 8.06 tok/s | 63.57 |
|
| 82 |
+
| prefill | 55.7 tok/s | 228.3 |
|
| 83 |
+
|
| 84 |
+
Load takes 0.48 s on CPU and 0.28 s on GPU,
|
| 85 |
+
the latter including the host-to-device upload.
|
| 86 |
+
|
| 87 |
+
|
| 88 |
117,530 paired tokens of FineWeb-Edu in **229 independent
|
| 89 |
512-token windows**. Both models score identical tokens and are
|
| 90 |
compared per token, which cuts the standard error 10.3x versus
|