Maggio33 commited on
Commit
8e55eb1
Β·
verified Β·
1 Parent(s): 73c44cb

Card: 64M 500k-540k, 128M 340k/360k, v1 128M 360k reference

Browse files
Files changed (1) hide show
  1. README.md +7 -1
README.md CHANGED
@@ -188,6 +188,9 @@ Two final models are training now with a recipe fixed before the runs started. N
188
  | 64M final (14Γ—576) | 440,000 (57.9%) | 46.17 | 75.74 | 2.3780 | 75.15 | training in progress; not a result |
189
  | 64M final (14Γ—576) | 460,000 (60.5%) | 46.30 | 75.47 | 2.3759 | 75.11 | training in progress; not a result |
190
  | 64M final (14Γ—576) | 480,000 (63.2%) | 46.72 | 75.69 | 2.3734 | 75.33 | training in progress; not a result |
 
 
 
191
  | 128M final (16Γ—768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
192
  | 128M final (16Γ—768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
193
  | 128M final (16Γ—768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
@@ -204,6 +207,8 @@ Two final models are training now with a recipe fixed before the runs started. N
204
  | 128M final (16Γ—768) | 280,000 (36.8%) | 50.17 | 78.74 | 2.3145 | 75.63 | training in progress; not a result |
205
  | 128M final (16Γ—768) | 300,000 (39.5%) | 50.46 | 78.07 | 2.3108 | 75.51 | training in progress; not a result |
206
  | 128M final (16Γ—768) | 320,000 (42.1%) | 50.46 | 77.79 | 2.3026 | 75.44 | training in progress; not a result |
 
 
207
 
208
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
209
 
@@ -223,9 +228,10 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
223
  | 128M v1 (`v1_128m/`) | 240,000 | 50.00 | 76.67 | 2.3094 | 74.89 |
224
  | 128M v1 (`v1_128m/`) | 280,000 | 51.14 | 77.10 | 2.2950 | 75.46 |
225
  | 128M v1 (`v1_128m/`) | 320,000 | 52.06 | 77.24 | 2.2812 | 75.85 |
 
226
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
227
 
228
- At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy βˆ’0.25, BLiMP βˆ’1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy βˆ’1.90, BLiMP βˆ’0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. At step 240,000 the final 128M run is 0.08 eff behind its v1 (74.81 vs 74.89: ARC-Easy βˆ’1.09, BLiMP +1.00, byte_ppl +0.0193), with v1's learning rate at about 41% of peak versus 80%. At step 400,000, where v1 64M ended, the final 64M run is 0.57 eff behind v1's result (75.24 vs 75.81: ARC-Easy βˆ’0.97, BLiMP βˆ’0.52, byte_ppl +0.0212); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. At step 280,000 the final 128M run is +0.17 eff ahead of its v1 (75.63 vs 75.46: ARC-Easy βˆ’0.97, BLiMP +1.64, byte_ppl +0.0195), with v1's learning rate at about 29% of peak versus 73%. At step 320,000 it is 0.41 eff behind (75.44 vs 75.85: ARC-Easy βˆ’1.60, BLiMP +0.55, byte_ppl +0.0214), with v1's learning rate at about 19% of peak versus 66%. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
229
 
230
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
231
 
 
188
  | 64M final (14Γ—576) | 440,000 (57.9%) | 46.17 | 75.74 | 2.3780 | 75.15 | training in progress; not a result |
189
  | 64M final (14Γ—576) | 460,000 (60.5%) | 46.30 | 75.47 | 2.3759 | 75.11 | training in progress; not a result |
190
  | 64M final (14Γ—576) | 480,000 (63.2%) | 46.72 | 75.69 | 2.3734 | 75.33 | training in progress; not a result |
191
+ | 64M final (14Γ—576) | 500,000 (65.8%) | 47.05 | 75.55 | 2.3693 | 75.41 | training in progress; not a result |
192
+ | 64M final (14Γ—576) | 520,000 (68.4%) | 46.80 | 75.60 | 2.3674 | 75.35 | training in progress; not a result |
193
+ | 64M final (14Γ—576) | 540,000 (71.1%) | 46.46 | 75.83 | 2.3678 | 75.31 | training in progress; not a result |
194
  | 128M final (16Γ—768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
195
  | 128M final (16Γ—768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
196
  | 128M final (16Γ—768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
 
207
  | 128M final (16Γ—768) | 280,000 (36.8%) | 50.17 | 78.74 | 2.3145 | 75.63 | training in progress; not a result |
208
  | 128M final (16Γ—768) | 300,000 (39.5%) | 50.46 | 78.07 | 2.3108 | 75.51 | training in progress; not a result |
209
  | 128M final (16Γ—768) | 320,000 (42.1%) | 50.46 | 77.79 | 2.3026 | 75.44 | training in progress; not a result |
210
+ | 128M final (16Γ—768) | 340,000 (44.7%) | 50.13 | 78.32 | 2.3040 | 75.50 | training in progress; not a result |
211
+ | 128M final (16Γ—768) | 360,000 (47.4%) | 50.84 | 78.01 | 2.3011 | 75.64 | training in progress; not a result |
212
 
213
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
214
 
 
228
  | 128M v1 (`v1_128m/`) | 240,000 | 50.00 | 76.67 | 2.3094 | 74.89 |
229
  | 128M v1 (`v1_128m/`) | 280,000 | 51.14 | 77.10 | 2.2950 | 75.46 |
230
  | 128M v1 (`v1_128m/`) | 320,000 | 52.06 | 77.24 | 2.2812 | 75.85 |
231
+ | 128M v1 (`v1_128m/`) | 360,000 | 52.53 | 77.12 | 2.2731 | 75.99 |
232
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
233
 
234
+ At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy βˆ’0.25, BLiMP βˆ’1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy βˆ’1.90, BLiMP βˆ’0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. At step 240,000 the final 128M run is 0.08 eff behind its v1 (74.81 vs 74.89: ARC-Easy βˆ’1.09, BLiMP +1.00, byte_ppl +0.0193), with v1's learning rate at about 41% of peak versus 80%. At step 400,000, where v1 64M ended, the final 64M run is 0.57 eff behind v1's result (75.24 vs 75.81: ARC-Easy βˆ’0.97, BLiMP βˆ’0.52, byte_ppl +0.0212); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. At step 280,000 the final 128M run is +0.17 eff ahead of its v1 (75.63 vs 75.46: ARC-Easy βˆ’0.97, BLiMP +1.64, byte_ppl +0.0195), with v1's learning rate at about 29% of peak versus 73%. At step 320,000 it is 0.41 eff behind (75.44 vs 75.85: ARC-Easy βˆ’1.60, BLiMP +0.55, byte_ppl +0.0214), with v1's learning rate at about 19% of peak versus 66%. At step 360,000 it is 0.35 eff behind (75.64 vs 75.99: ARC-Easy βˆ’1.69, BLiMP +0.89, byte_ppl +0.0280), with v1's learning rate at about 12% of peak versus 59%. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
235
 
236
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
237