Card: 128M final @120k and v1 128M @120k same-step reference
Browse files
README.md
CHANGED
|
@@ -176,6 +176,7 @@ Two final models are training now with a recipe fixed before the runs started. N
|
|
| 176 |
| 128M final (16Γ768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
|
| 177 |
| 128M final (16Γ768) | 80,000 (10.5%) | 46.59 | 77.57 | 2.3959 | 73.83 | training in progress; not a result |
|
| 178 |
| 128M final (16Γ768) | 100,000 (13.2%) | 47.81 | 77.35 | 2.3813 | 74.20 | training in progress; not a result |
|
|
|
|
| 179 |
|
| 180 |
Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
|
| 181 |
|
|
@@ -187,9 +188,10 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
|
|
| 187 |
| 64M v1 (`v1_muon/`), end of run | 400,000 | 47.94 | 75.83 | 2.3718 | 75.81 |
|
| 188 |
| 128M v1 (`v1_128m/`) | 40,000 | 45.83 | 73.21 | 2.4577 | 71.95 |
|
| 189 |
| 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.04 |
|
|
|
|
| 190 |
| 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
|
| 191 |
|
| 192 |
-
At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl β0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP β1.72, byte_ppl β0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl β0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP β1.71, byte_ppl β0.0020). Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
|
| 193 |
|
| 194 |
## Data-attribution experiments (64M, 2026-09-25)
|
| 195 |
|
|
|
|
| 176 |
| 128M final (16Γ768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
|
| 177 |
| 128M final (16Γ768) | 80,000 (10.5%) | 46.59 | 77.57 | 2.3959 | 73.83 | training in progress; not a result |
|
| 178 |
| 128M final (16Γ768) | 100,000 (13.2%) | 47.81 | 77.35 | 2.3813 | 74.20 | training in progress; not a result |
|
| 179 |
+
| 128M final (16Γ768) | 120,000 (15.8%) | 48.82 | 77.51 | 2.3760 | 74.61 | training in progress; not a result |
|
| 180 |
|
| 181 |
Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
|
| 182 |
|
|
|
|
| 188 |
| 64M v1 (`v1_muon/`), end of run | 400,000 | 47.94 | 75.83 | 2.3718 | 75.81 |
|
| 189 |
| 128M v1 (`v1_128m/`) | 40,000 | 45.83 | 73.21 | 2.4577 | 71.95 |
|
| 190 |
| 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.04 |
|
| 191 |
+
| 128M v1 (`v1_128m/`), last checkpoint before its resume | 120,000 | 48.02 | 75.19 | 2.3628 | 73.59 |
|
| 192 |
| 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
|
| 193 |
|
| 194 |
+
At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl β0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP β1.72, byte_ppl β0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl β0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP β1.71, byte_ppl β0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
|
| 195 |
|
| 196 |
## Data-attribution experiments (64M, 2026-09-25)
|
| 197 |
|