Maggio33 commited on
Commit
4c6c541
Β·
verified Β·
1 Parent(s): cf5a3a2

Card: 64M 620k, 128M 400k, 400k comparison with v1 128M end of run

Browse files
Files changed (1) hide show
  1. README.md +3 -1
README.md CHANGED
@@ -194,6 +194,7 @@ Two final models are training now with a recipe fixed before the runs started. N
194
  | 64M final (14Γ—576) | 560,000 (73.7%) | 47.81 | 76.17 | 2.3648 | 75.90 | training in progress; not a result |
195
  | 64M final (14Γ—576) | 580,000 (76.3%) | 47.47 | 76.38 | 2.3604 | 75.87 | training in progress; not a result |
196
  | 64M final (14Γ—576) | 600,000 (78.9%) | 47.56 | 76.11 | 2.3561 | 75.81 | training in progress; not a result |
 
197
  | 128M final (16Γ—768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
198
  | 128M final (16Γ—768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
199
  | 128M final (16Γ—768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
@@ -213,6 +214,7 @@ Two final models are training now with a recipe fixed before the runs started. N
213
  | 128M final (16Γ—768) | 340,000 (44.7%) | 50.13 | 78.32 | 2.3040 | 75.50 | training in progress; not a result |
214
  | 128M final (16Γ—768) | 360,000 (47.4%) | 50.84 | 78.01 | 2.3011 | 75.64 | training in progress; not a result |
215
  | 128M final (16Γ—768) | 380,000 (50.0%) | 51.18 | 78.86 | 2.2994 | 76.05 | training in progress; not a result |
 
216
 
217
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
218
 
@@ -235,7 +237,7 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
235
  | 128M v1 (`v1_128m/`) | 360,000 | 52.53 | 77.12 | 2.2731 | 75.99 |
236
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
237
 
238
- At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy βˆ’0.25, BLiMP βˆ’1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy βˆ’1.90, BLiMP βˆ’0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. At step 240,000 the final 128M run is 0.08 eff behind its v1 (74.81 vs 74.89: ARC-Easy βˆ’1.09, BLiMP +1.00, byte_ppl +0.0193), with v1's learning rate at about 41% of peak versus 80%. At step 400,000, where v1 64M ended, the final 64M run is 0.57 eff behind v1's result (75.24 vs 75.81: ARC-Easy βˆ’0.97, BLiMP βˆ’0.52, byte_ppl +0.0212); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. At step 280,000 the final 128M run is +0.17 eff ahead of its v1 (75.63 vs 75.46: ARC-Easy βˆ’0.97, BLiMP +1.64, byte_ppl +0.0195), with v1's learning rate at about 29% of peak versus 73%. At step 320,000 it is 0.41 eff behind (75.44 vs 75.85: ARC-Easy βˆ’1.60, BLiMP +0.55, byte_ppl +0.0214), with v1's learning rate at about 19% of peak versus 66%. At step 360,000 it is 0.35 eff behind (75.64 vs 75.99: ARC-Easy βˆ’1.69, BLiMP +0.89, byte_ppl +0.0280), with v1's learning rate at about 12% of peak versus 59%. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
239
 
240
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
241
 
 
194
  | 64M final (14Γ—576) | 560,000 (73.7%) | 47.81 | 76.17 | 2.3648 | 75.90 | training in progress; not a result |
195
  | 64M final (14Γ—576) | 580,000 (76.3%) | 47.47 | 76.38 | 2.3604 | 75.87 | training in progress; not a result |
196
  | 64M final (14Γ—576) | 600,000 (78.9%) | 47.56 | 76.11 | 2.3561 | 75.81 | training in progress; not a result |
197
+ | 64M final (14Γ—576) | 620,000 (81.6%) | 48.78 | 75.96 | 2.3589 | 76.18 | training in progress; not a result |
198
  | 128M final (16Γ—768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
199
  | 128M final (16Γ—768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
200
  | 128M final (16Γ—768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
 
214
  | 128M final (16Γ—768) | 340,000 (44.7%) | 50.13 | 78.32 | 2.3040 | 75.50 | training in progress; not a result |
215
  | 128M final (16Γ—768) | 360,000 (47.4%) | 50.84 | 78.01 | 2.3011 | 75.64 | training in progress; not a result |
216
  | 128M final (16Γ—768) | 380,000 (50.0%) | 51.18 | 78.86 | 2.2994 | 76.05 | training in progress; not a result |
217
+ | 128M final (16Γ—768) | 400,000 (52.6%) | 51.73 | 79.38 | 2.2966 | 76.42 | training in progress; not a result |
218
 
219
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
220
 
 
237
  | 128M v1 (`v1_128m/`) | 360,000 | 52.53 | 77.12 | 2.2731 | 75.99 |
238
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
239
 
240
+ At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy βˆ’0.25, BLiMP βˆ’1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy βˆ’1.90, BLiMP βˆ’0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. At step 240,000 the final 128M run is 0.08 eff behind its v1 (74.81 vs 74.89: ARC-Easy βˆ’1.09, BLiMP +1.00, byte_ppl +0.0193), with v1's learning rate at about 41% of peak versus 80%. At step 400,000, where v1 64M ended, the final 64M run is 0.57 eff behind v1's result (75.24 vs 75.81: ARC-Easy βˆ’0.97, BLiMP βˆ’0.52, byte_ppl +0.0212); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. At step 280,000 the final 128M run is +0.17 eff ahead of its v1 (75.63 vs 75.46: ARC-Easy βˆ’0.97, BLiMP +1.64, byte_ppl +0.0195), with v1's learning rate at about 29% of peak versus 73%. At step 320,000 it is 0.41 eff behind (75.44 vs 75.85: ARC-Easy βˆ’1.60, BLiMP +0.55, byte_ppl +0.0214), with v1's learning rate at about 19% of peak versus 66%. At step 360,000 it is 0.35 eff behind (75.64 vs 75.99: ARC-Easy βˆ’1.69, BLiMP +0.89, byte_ppl +0.0280), with v1's learning rate at about 12% of peak versus 59%. At step 400,000, where v1 128M ended, the final 128M run is 0.58 eff ahead of v1's result (76.42 vs 75.84: ARC-Easy βˆ’0.21, BLiMP +2.12, byte_ppl +0.0249); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
241
 
242
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
243