Card: 64M 620k, 128M 400k, 400k comparison with v1 128M end of run
Browse files
README.md
CHANGED
|
@@ -194,6 +194,7 @@ Two final models are training now with a recipe fixed before the runs started. N
|
|
| 194 |
| 64M final (14Γ576) | 560,000 (73.7%) | 47.81 | 76.17 | 2.3648 | 75.90 | training in progress; not a result |
|
| 195 |
| 64M final (14Γ576) | 580,000 (76.3%) | 47.47 | 76.38 | 2.3604 | 75.87 | training in progress; not a result |
|
| 196 |
| 64M final (14Γ576) | 600,000 (78.9%) | 47.56 | 76.11 | 2.3561 | 75.81 | training in progress; not a result |
|
|
|
|
| 197 |
| 128M final (16Γ768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
|
| 198 |
| 128M final (16Γ768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
|
| 199 |
| 128M final (16Γ768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
|
|
@@ -213,6 +214,7 @@ Two final models are training now with a recipe fixed before the runs started. N
|
|
| 213 |
| 128M final (16Γ768) | 340,000 (44.7%) | 50.13 | 78.32 | 2.3040 | 75.50 | training in progress; not a result |
|
| 214 |
| 128M final (16Γ768) | 360,000 (47.4%) | 50.84 | 78.01 | 2.3011 | 75.64 | training in progress; not a result |
|
| 215 |
| 128M final (16Γ768) | 380,000 (50.0%) | 51.18 | 78.86 | 2.2994 | 76.05 | training in progress; not a result |
|
|
|
|
| 216 |
|
| 217 |
Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
|
| 218 |
|
|
@@ -235,7 +237,7 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
|
|
| 235 |
| 128M v1 (`v1_128m/`) | 360,000 | 52.53 | 77.12 | 2.2731 | 75.99 |
|
| 236 |
| 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
|
| 237 |
|
| 238 |
-
At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl β0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP β1.72, byte_ppl β0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl β0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP β1.71, byte_ppl β0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy β1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy β0.25, BLiMP β1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy β1.90, BLiMP β0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. At step 240,000 the final 128M run is 0.08 eff behind its v1 (74.81 vs 74.89: ARC-Easy β1.09, BLiMP +1.00, byte_ppl +0.0193), with v1's learning rate at about 41% of peak versus 80%. At step 400,000, where v1 64M ended, the final 64M run is 0.57 eff behind v1's result (75.24 vs 75.81: ARC-Easy β0.97, BLiMP β0.52, byte_ppl +0.0212); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. At step 280,000 the final 128M run is +0.17 eff ahead of its v1 (75.63 vs 75.46: ARC-Easy β0.97, BLiMP +1.64, byte_ppl +0.0195), with v1's learning rate at about 29% of peak versus 73%. At step 320,000 it is 0.41 eff behind (75.44 vs 75.85: ARC-Easy β1.60, BLiMP +0.55, byte_ppl +0.0214), with v1's learning rate at about 19% of peak versus 66%. At step 360,000 it is 0.35 eff behind (75.64 vs 75.99: ARC-Easy β1.69, BLiMP +0.89, byte_ppl +0.0280), with v1's learning rate at about 12% of peak versus 59%. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
|
| 239 |
|
| 240 |
**Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
|
| 241 |
|
|
|
|
| 194 |
| 64M final (14Γ576) | 560,000 (73.7%) | 47.81 | 76.17 | 2.3648 | 75.90 | training in progress; not a result |
|
| 195 |
| 64M final (14Γ576) | 580,000 (76.3%) | 47.47 | 76.38 | 2.3604 | 75.87 | training in progress; not a result |
|
| 196 |
| 64M final (14Γ576) | 600,000 (78.9%) | 47.56 | 76.11 | 2.3561 | 75.81 | training in progress; not a result |
|
| 197 |
+
| 64M final (14Γ576) | 620,000 (81.6%) | 48.78 | 75.96 | 2.3589 | 76.18 | training in progress; not a result |
|
| 198 |
| 128M final (16Γ768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
|
| 199 |
| 128M final (16Γ768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
|
| 200 |
| 128M final (16Γ768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
|
|
|
|
| 214 |
| 128M final (16Γ768) | 340,000 (44.7%) | 50.13 | 78.32 | 2.3040 | 75.50 | training in progress; not a result |
|
| 215 |
| 128M final (16Γ768) | 360,000 (47.4%) | 50.84 | 78.01 | 2.3011 | 75.64 | training in progress; not a result |
|
| 216 |
| 128M final (16Γ768) | 380,000 (50.0%) | 51.18 | 78.86 | 2.2994 | 76.05 | training in progress; not a result |
|
| 217 |
+
| 128M final (16Γ768) | 400,000 (52.6%) | 51.73 | 79.38 | 2.2966 | 76.42 | training in progress; not a result |
|
| 218 |
|
| 219 |
Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
|
| 220 |
|
|
|
|
| 237 |
| 128M v1 (`v1_128m/`) | 360,000 | 52.53 | 77.12 | 2.2731 | 75.99 |
|
| 238 |
| 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
|
| 239 |
|
| 240 |
+
At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl β0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP β1.72, byte_ppl β0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl β0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP β1.71, byte_ppl β0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy β1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy β0.25, BLiMP β1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy β1.90, BLiMP β0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. At step 240,000 the final 128M run is 0.08 eff behind its v1 (74.81 vs 74.89: ARC-Easy β1.09, BLiMP +1.00, byte_ppl +0.0193), with v1's learning rate at about 41% of peak versus 80%. At step 400,000, where v1 64M ended, the final 64M run is 0.57 eff behind v1's result (75.24 vs 75.81: ARC-Easy β0.97, BLiMP β0.52, byte_ppl +0.0212); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. At step 280,000 the final 128M run is +0.17 eff ahead of its v1 (75.63 vs 75.46: ARC-Easy β0.97, BLiMP +1.64, byte_ppl +0.0195), with v1's learning rate at about 29% of peak versus 73%. At step 320,000 it is 0.41 eff behind (75.44 vs 75.85: ARC-Easy β1.60, BLiMP +0.55, byte_ppl +0.0214), with v1's learning rate at about 19% of peak versus 66%. At step 360,000 it is 0.35 eff behind (75.64 vs 75.99: ARC-Easy β1.69, BLiMP +0.89, byte_ppl +0.0280), with v1's learning rate at about 12% of peak versus 59%. At step 400,000, where v1 128M ended, the final 128M run is 0.58 eff ahead of v1's result (76.42 vs 75.84: ARC-Easy β0.21, BLiMP +2.12, byte_ppl +0.0249); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
|
| 241 |
|
| 242 |
**Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
|
| 243 |
|