Maggio33 commited on
Commit
aaaebee
Β·
verified Β·
1 Parent(s): fb0ca1e

Card: 64M 340k/360k, 128M 220k, v1 360k reference

Browse files
Files changed (1) hide show
  1. README.md +5 -1
README.md CHANGED
@@ -180,6 +180,8 @@ Two final models are training now with a recipe fixed before the runs started. N
180
  | 64M final (14Γ—576) | 280,000 (36.8%) | 46.51 | 74.48 | 2.4016 | 74.77 | training in progress; not a result |
181
  | 64M final (14Γ—576) | 300,000 (39.5%) | 46.59 | 74.41 | 2.4003 | 74.78 | training in progress; not a result |
182
  | 64M final (14Γ—576) | 320,000 (42.1%), first after the resume | 46.97 | 74.56 | 2.3955 | 74.97 | training in progress; not a result |
 
 
183
  | 128M final (16Γ—768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
184
  | 128M final (16Γ—768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
185
  | 128M final (16Γ—768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
@@ -190,6 +192,7 @@ Two final models are training now with a recipe fixed before the runs started. N
190
  | 128M final (16Γ—768) | 160,000 (21.1%) | 49.62 | 77.14 | 2.3523 | 74.81 | training in progress; not a result |
191
  | 128M final (16Γ—768) | 180,000 (23.7%) | 49.41 | 77.65 | 2.3455 | 74.93 | training in progress; not a result |
192
  | 128M final (16Γ—768) | 200,000 (26.3%) | 48.36 | 78.65 | 2.3406 | 74.92 | training in progress; not a result |
 
193
 
194
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
195
 
@@ -199,6 +202,7 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
199
  | 64M v1 (`v1_muon/`) | 80,000 | 43.48 | 75.24 | 2.4922 | 73.76 |
200
  | 64M v1 (`v1_muon/`) | 160,000 | 44.99 | 75.52 | 2.4436 | 74.50 |
201
  | 64M v1 (`v1_muon/`) | 320,000 | 47.22 | 75.89 | 2.3775 | 75.57 |
 
202
  | 64M v1 (`v1_muon/`), end of run | 400,000 | 47.94 | 75.83 | 2.3718 | 75.81 |
203
  | 128M v1 (`v1_128m/`) | 40,000 | 45.83 | 73.21 | 2.4577 | 71.95 |
204
  | 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.04 |
@@ -207,7 +211,7 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
207
  | 128M v1 (`v1_128m/`) | 200,000 | 49.62 | 76.78 | 2.3344 | 74.74 |
208
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
209
 
210
- At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy βˆ’0.25, BLiMP βˆ’1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
211
 
212
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
213
 
 
180
  | 64M final (14Γ—576) | 280,000 (36.8%) | 46.51 | 74.48 | 2.4016 | 74.77 | training in progress; not a result |
181
  | 64M final (14Γ—576) | 300,000 (39.5%) | 46.59 | 74.41 | 2.4003 | 74.78 | training in progress; not a result |
182
  | 64M final (14Γ—576) | 320,000 (42.1%), first after the resume | 46.97 | 74.56 | 2.3955 | 74.97 | training in progress; not a result |
183
+ | 64M final (14Γ—576) | 340,000 (44.7%) | 46.80 | 75.05 | 2.3987 | 75.08 | training in progress; not a result |
184
+ | 64M final (14Γ—576) | 360,000 (47.4%) | 45.79 | 75.26 | 2.3917 | 74.82 | training in progress; not a result |
185
  | 128M final (16Γ—768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
186
  | 128M final (16Γ—768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
187
  | 128M final (16Γ—768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
 
192
  | 128M final (16Γ—768) | 160,000 (21.1%) | 49.62 | 77.14 | 2.3523 | 74.81 | training in progress; not a result |
193
  | 128M final (16Γ—768) | 180,000 (23.7%) | 49.41 | 77.65 | 2.3455 | 74.93 | training in progress; not a result |
194
  | 128M final (16Γ—768) | 200,000 (26.3%) | 48.36 | 78.65 | 2.3406 | 74.92 | training in progress; not a result |
195
+ | 128M final (16Γ—768) | 220,000 (28.9%), first after the resume | 49.37 | 78.53 | 2.3275 | 75.26 | training in progress; not a result |
196
 
197
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
198
 
 
202
  | 64M v1 (`v1_muon/`) | 80,000 | 43.48 | 75.24 | 2.4922 | 73.76 |
203
  | 64M v1 (`v1_muon/`) | 160,000 | 44.99 | 75.52 | 2.4436 | 74.50 |
204
  | 64M v1 (`v1_muon/`) | 320,000 | 47.22 | 75.89 | 2.3775 | 75.57 |
205
+ | 64M v1 (`v1_muon/`) | 360,000 | 47.69 | 76.15 | 2.3746 | 75.83 |
206
  | 64M v1 (`v1_muon/`), end of run | 400,000 | 47.94 | 75.83 | 2.3718 | 75.81 |
207
  | 128M v1 (`v1_128m/`) | 40,000 | 45.83 | 73.21 | 2.4577 | 71.95 |
208
  | 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.04 |
 
211
  | 128M v1 (`v1_128m/`) | 200,000 | 49.62 | 76.78 | 2.3344 | 74.74 |
212
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
213
 
214
+ At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy βˆ’0.25, BLiMP βˆ’1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy βˆ’1.90, BLiMP βˆ’0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
215
 
216
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
217