Maggio33 commited on
Commit
54f12c1
Β·
verified Β·
1 Parent(s): 7c657ee

Card: 64M 440k-480k, 128M 300k/320k, v1 128M 320k reference

Browse files
Files changed (1) hide show
  1. README.md +7 -1
README.md CHANGED
@@ -185,6 +185,9 @@ Two final models are training now with a recipe fixed before the runs started. N
185
  | 64M final (14Γ—576) | 380,000 (50.0%) | 46.72 | 75.14 | 2.3957 | 75.09 | training in progress; not a result |
186
  | 64M final (14Γ—576) | 400,000 (52.6%) | 46.97 | 75.31 | 2.3930 | 75.24 | training in progress; not a result |
187
  | 64M final (14Γ—576) | 420,000 (55.3%) | 47.56 | 75.46 | 2.3865 | 75.51 | training in progress; not a result |
 
 
 
188
  | 128M final (16Γ—768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
189
  | 128M final (16Γ—768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
190
  | 128M final (16Γ—768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
@@ -199,6 +202,8 @@ Two final models are training now with a recipe fixed before the runs started. N
199
  | 128M final (16Γ—768) | 240,000 (31.6%) | 48.91 | 77.67 | 2.3287 | 74.81 | training in progress; not a result |
200
  | 128M final (16Γ—768) | 260,000 (34.2%) | 49.79 | 78.53 | 2.3150 | 75.43 | training in progress; not a result |
201
  | 128M final (16Γ—768) | 280,000 (36.8%) | 50.17 | 78.74 | 2.3145 | 75.63 | training in progress; not a result |
 
 
202
 
203
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
204
 
@@ -217,9 +222,10 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
217
  | 128M v1 (`v1_128m/`) | 200,000 | 49.62 | 76.78 | 2.3344 | 74.74 |
218
  | 128M v1 (`v1_128m/`) | 240,000 | 50.00 | 76.67 | 2.3094 | 74.89 |
219
  | 128M v1 (`v1_128m/`) | 280,000 | 51.14 | 77.10 | 2.2950 | 75.46 |
 
220
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
221
 
222
- At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy βˆ’0.25, BLiMP βˆ’1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy βˆ’1.90, BLiMP βˆ’0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. At step 240,000 the final 128M run is 0.08 eff behind its v1 (74.81 vs 74.89: ARC-Easy βˆ’1.09, BLiMP +1.00, byte_ppl +0.0193), with v1's learning rate at about 41% of peak versus 80%. At step 400,000, where v1 64M ended, the final 64M run is 0.57 eff behind v1's result (75.24 vs 75.81: ARC-Easy βˆ’0.97, BLiMP βˆ’0.52, byte_ppl +0.0212); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. At step 280,000 the final 128M run is +0.17 eff ahead of its v1 (75.63 vs 75.46: ARC-Easy βˆ’0.97, BLiMP +1.64, byte_ppl +0.0195), with v1's learning rate at about 29% of peak versus 73%. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
223
 
224
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
225
 
 
185
  | 64M final (14Γ—576) | 380,000 (50.0%) | 46.72 | 75.14 | 2.3957 | 75.09 | training in progress; not a result |
186
  | 64M final (14Γ—576) | 400,000 (52.6%) | 46.97 | 75.31 | 2.3930 | 75.24 | training in progress; not a result |
187
  | 64M final (14Γ—576) | 420,000 (55.3%) | 47.56 | 75.46 | 2.3865 | 75.51 | training in progress; not a result |
188
+ | 64M final (14Γ—576) | 440,000 (57.9%) | 46.17 | 75.74 | 2.3780 | 75.15 | training in progress; not a result |
189
+ | 64M final (14Γ—576) | 460,000 (60.5%) | 46.30 | 75.47 | 2.3759 | 75.11 | training in progress; not a result |
190
+ | 64M final (14Γ—576) | 480,000 (63.2%) | 46.72 | 75.69 | 2.3734 | 75.33 | training in progress; not a result |
191
  | 128M final (16Γ—768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
192
  | 128M final (16Γ—768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
193
  | 128M final (16Γ—768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
 
202
  | 128M final (16Γ—768) | 240,000 (31.6%) | 48.91 | 77.67 | 2.3287 | 74.81 | training in progress; not a result |
203
  | 128M final (16Γ—768) | 260,000 (34.2%) | 49.79 | 78.53 | 2.3150 | 75.43 | training in progress; not a result |
204
  | 128M final (16Γ—768) | 280,000 (36.8%) | 50.17 | 78.74 | 2.3145 | 75.63 | training in progress; not a result |
205
+ | 128M final (16Γ—768) | 300,000 (39.5%) | 50.46 | 78.07 | 2.3108 | 75.51 | training in progress; not a result |
206
+ | 128M final (16Γ—768) | 320,000 (42.1%) | 50.46 | 77.79 | 2.3026 | 75.44 | training in progress; not a result |
207
 
208
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
209
 
 
222
  | 128M v1 (`v1_128m/`) | 200,000 | 49.62 | 76.78 | 2.3344 | 74.74 |
223
  | 128M v1 (`v1_128m/`) | 240,000 | 50.00 | 76.67 | 2.3094 | 74.89 |
224
  | 128M v1 (`v1_128m/`) | 280,000 | 51.14 | 77.10 | 2.2950 | 75.46 |
225
+ | 128M v1 (`v1_128m/`) | 320,000 | 52.06 | 77.24 | 2.2812 | 75.85 |
226
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
227
 
228
+ At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy βˆ’0.25, BLiMP βˆ’1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy βˆ’1.90, BLiMP βˆ’0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. At step 240,000 the final 128M run is 0.08 eff behind its v1 (74.81 vs 74.89: ARC-Easy βˆ’1.09, BLiMP +1.00, byte_ppl +0.0193), with v1's learning rate at about 41% of peak versus 80%. At step 400,000, where v1 64M ended, the final 64M run is 0.57 eff behind v1's result (75.24 vs 75.81: ARC-Easy βˆ’0.97, BLiMP βˆ’0.52, byte_ppl +0.0212); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. At step 280,000 the final 128M run is +0.17 eff ahead of its v1 (75.63 vs 75.46: ARC-Easy βˆ’0.97, BLiMP +1.64, byte_ppl +0.0195), with v1's learning rate at about 29% of peak versus 73%. At step 320,000 it is 0.41 eff behind (75.44 vs 75.85: ARC-Easy βˆ’1.60, BLiMP +0.55, byte_ppl +0.0214), with v1's learning rate at about 19% of peak versus 66%. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
229
 
230
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
231