Maggio33 commited on
Commit
e6ed1be
·
verified ·
1 Parent(s): 6e221fd

README: progress table (64M 100k, 128M 60k/80k), 128M same-step comparison at 80k

Browse files
Files changed (1) hide show
  1. README.md +4 -1
README.md CHANGED
@@ -167,8 +167,11 @@ Two final models are training now with a recipe fixed before the runs started. N
167
  | 64M final (14×576) | 40,000 (5.3%) | 43.43 | 73.46 | 2.5388 | 73.01 | training in progress; not a result |
168
  | 64M final (14×576) | 60,000 (7.9%) | 43.77 | 73.41 | 2.4992 | 73.21 | training in progress; not a result |
169
  | 64M final (14×576) | 80,000 (10.5%) | 45.08 | 73.52 | 2.4777 | 73.75 | training in progress; not a result |
 
170
  | 128M final (16×768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
171
  | 128M final (16×768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
 
 
172
 
173
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
174
 
@@ -181,7 +184,7 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
181
  | 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.03 |
182
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
183
 
184
- At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl −0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP −1.72, byte_ppl −0.0145). Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps), so from roughly 100k steps on the same-step comparison increasingly favours v1.
185
 
186
  ## Data-attribution experiments (64M, 2026-09-25)
187
 
 
167
  | 64M final (14×576) | 40,000 (5.3%) | 43.43 | 73.46 | 2.5388 | 73.01 | training in progress; not a result |
168
  | 64M final (14×576) | 60,000 (7.9%) | 43.77 | 73.41 | 2.4992 | 73.21 | training in progress; not a result |
169
  | 64M final (14×576) | 80,000 (10.5%) | 45.08 | 73.52 | 2.4777 | 73.75 | training in progress; not a result |
170
+ | 64M final (14×576) | 100,000 (13.2%) | 45.50 | 74.49 | 2.4733 | 74.24 | training in progress; not a result |
171
  | 128M final (16×768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
172
  | 128M final (16×768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
173
+ | 128M final (16×768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
174
+ | 128M final (16×768) | 80,000 (10.5%) | 46.59 | 77.57 | 2.3959 | 73.83 | training in progress; not a result |
175
 
176
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
177
 
 
184
  | 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.03 |
185
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
186
 
187
+ At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl −0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP −1.72, byte_ppl −0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.03: ARC-Easy +0.21, BLiMP +2.04, byte_ppl −0.0137). Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps), so from roughly 100k steps on the same-step comparison increasingly favours v1.
188
 
189
  ## Data-attribution experiments (64M, 2026-09-25)
190