README: progress table @80k, v1 128M @80k reference, same-step comparison at 80k
Browse files
README.md
CHANGED
|
@@ -166,6 +166,7 @@ Two final models are training now with a recipe fixed before the runs started. N
|
|
| 166 |
| 64M final (14×576) | 20,000 (2.6%) | 40.61 | 72.83 | 2.6070 | 71.66 | training in progress; not a result |
|
| 167 |
| 64M final (14×576) | 40,000 (5.3%) | 43.43 | 73.46 | 2.5388 | 73.01 | training in progress; not a result |
|
| 168 |
| 64M final (14×576) | 60,000 (7.9%) | 43.77 | 73.41 | 2.4992 | 73.21 | training in progress; not a result |
|
|
|
|
| 169 |
| 128M final (16×768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
|
| 170 |
| 128M final (16×768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
|
| 171 |
|
|
@@ -177,9 +178,10 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
|
|
| 177 |
| 64M v1 (`v1_muon/`) | 80,000 | 43.48 | 75.24 | 2.4922 | 73.76 |
|
| 178 |
| 64M v1 (`v1_muon/`), end of run | 400,000 | 47.94 | 75.83 | 2.3718 | 75.81 |
|
| 179 |
| 128M v1 (`v1_128m/`) | 40,000 | 45.83 | 73.21 | 2.4577 | 71.95 |
|
|
|
|
| 180 |
| 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
|
| 181 |
|
| 182 |
-
At step 40,000 the final 64M run
|
| 183 |
|
| 184 |
## Data-attribution experiments (64M, 2026-09-25)
|
| 185 |
|
|
|
|
| 166 |
| 64M final (14×576) | 20,000 (2.6%) | 40.61 | 72.83 | 2.6070 | 71.66 | training in progress; not a result |
|
| 167 |
| 64M final (14×576) | 40,000 (5.3%) | 43.43 | 73.46 | 2.5388 | 73.01 | training in progress; not a result |
|
| 168 |
| 64M final (14×576) | 60,000 (7.9%) | 43.77 | 73.41 | 2.4992 | 73.21 | training in progress; not a result |
|
| 169 |
+
| 64M final (14×576) | 80,000 (10.5%) | 45.08 | 73.52 | 2.4777 | 73.75 | training in progress; not a result |
|
| 170 |
| 128M final (16×768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
|
| 171 |
| 128M final (16×768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
|
| 172 |
|
|
|
|
| 178 |
| 64M v1 (`v1_muon/`) | 80,000 | 43.48 | 75.24 | 2.4922 | 73.76 |
|
| 179 |
| 64M v1 (`v1_muon/`), end of run | 400,000 | 47.94 | 75.83 | 2.3718 | 75.81 |
|
| 180 |
| 128M v1 (`v1_128m/`) | 40,000 | 45.83 | 73.21 | 2.4577 | 71.95 |
|
| 181 |
+
| 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.03 |
|
| 182 |
| 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
|
| 183 |
|
| 184 |
+
At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl −0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP −1.72, byte_ppl −0.0145). Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps), so from roughly 100k steps on the same-step comparison increasingly favours v1.
|
| 185 |
|
| 186 |
## Data-attribution experiments (64M, 2026-09-25)
|
| 187 |
|