Card: 64M 500k-540k, 128M 340k/360k, v1 128M 360k reference
Browse files
README.md
CHANGED
|
@@ -188,6 +188,9 @@ Two final models are training now with a recipe fixed before the runs started. N
|
|
| 188 |
| 64M final (14Γ576) | 440,000 (57.9%) | 46.17 | 75.74 | 2.3780 | 75.15 | training in progress; not a result |
|
| 189 |
| 64M final (14Γ576) | 460,000 (60.5%) | 46.30 | 75.47 | 2.3759 | 75.11 | training in progress; not a result |
|
| 190 |
| 64M final (14Γ576) | 480,000 (63.2%) | 46.72 | 75.69 | 2.3734 | 75.33 | training in progress; not a result |
|
|
|
|
|
|
|
|
|
|
| 191 |
| 128M final (16Γ768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
|
| 192 |
| 128M final (16Γ768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
|
| 193 |
| 128M final (16Γ768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
|
|
@@ -204,6 +207,8 @@ Two final models are training now with a recipe fixed before the runs started. N
|
|
| 204 |
| 128M final (16Γ768) | 280,000 (36.8%) | 50.17 | 78.74 | 2.3145 | 75.63 | training in progress; not a result |
|
| 205 |
| 128M final (16Γ768) | 300,000 (39.5%) | 50.46 | 78.07 | 2.3108 | 75.51 | training in progress; not a result |
|
| 206 |
| 128M final (16Γ768) | 320,000 (42.1%) | 50.46 | 77.79 | 2.3026 | 75.44 | training in progress; not a result |
|
|
|
|
|
|
|
| 207 |
|
| 208 |
Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
|
| 209 |
|
|
@@ -223,9 +228,10 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
|
|
| 223 |
| 128M v1 (`v1_128m/`) | 240,000 | 50.00 | 76.67 | 2.3094 | 74.89 |
|
| 224 |
| 128M v1 (`v1_128m/`) | 280,000 | 51.14 | 77.10 | 2.2950 | 75.46 |
|
| 225 |
| 128M v1 (`v1_128m/`) | 320,000 | 52.06 | 77.24 | 2.2812 | 75.85 |
|
|
|
|
| 226 |
| 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
|
| 227 |
|
| 228 |
-
At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl β0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP β1.72, byte_ppl β0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl β0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP β1.71, byte_ppl β0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy β1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy β0.25, BLiMP β1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy β1.90, BLiMP β0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. At step 240,000 the final 128M run is 0.08 eff behind its v1 (74.81 vs 74.89: ARC-Easy β1.09, BLiMP +1.00, byte_ppl +0.0193), with v1's learning rate at about 41% of peak versus 80%. At step 400,000, where v1 64M ended, the final 64M run is 0.57 eff behind v1's result (75.24 vs 75.81: ARC-Easy β0.97, BLiMP β0.52, byte_ppl +0.0212); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. At step 280,000 the final 128M run is +0.17 eff ahead of its v1 (75.63 vs 75.46: ARC-Easy β0.97, BLiMP +1.64, byte_ppl +0.0195), with v1's learning rate at about 29% of peak versus 73%. At step 320,000 it is 0.41 eff behind (75.44 vs 75.85: ARC-Easy β1.60, BLiMP +0.55, byte_ppl +0.0214), with v1's learning rate at about 19% of peak versus 66%. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
|
| 229 |
|
| 230 |
**Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
|
| 231 |
|
|
|
|
| 188 |
| 64M final (14Γ576) | 440,000 (57.9%) | 46.17 | 75.74 | 2.3780 | 75.15 | training in progress; not a result |
|
| 189 |
| 64M final (14Γ576) | 460,000 (60.5%) | 46.30 | 75.47 | 2.3759 | 75.11 | training in progress; not a result |
|
| 190 |
| 64M final (14Γ576) | 480,000 (63.2%) | 46.72 | 75.69 | 2.3734 | 75.33 | training in progress; not a result |
|
| 191 |
+
| 64M final (14Γ576) | 500,000 (65.8%) | 47.05 | 75.55 | 2.3693 | 75.41 | training in progress; not a result |
|
| 192 |
+
| 64M final (14Γ576) | 520,000 (68.4%) | 46.80 | 75.60 | 2.3674 | 75.35 | training in progress; not a result |
|
| 193 |
+
| 64M final (14Γ576) | 540,000 (71.1%) | 46.46 | 75.83 | 2.3678 | 75.31 | training in progress; not a result |
|
| 194 |
| 128M final (16Γ768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
|
| 195 |
| 128M final (16Γ768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
|
| 196 |
| 128M final (16Γ768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
|
|
|
|
| 207 |
| 128M final (16Γ768) | 280,000 (36.8%) | 50.17 | 78.74 | 2.3145 | 75.63 | training in progress; not a result |
|
| 208 |
| 128M final (16Γ768) | 300,000 (39.5%) | 50.46 | 78.07 | 2.3108 | 75.51 | training in progress; not a result |
|
| 209 |
| 128M final (16Γ768) | 320,000 (42.1%) | 50.46 | 77.79 | 2.3026 | 75.44 | training in progress; not a result |
|
| 210 |
+
| 128M final (16Γ768) | 340,000 (44.7%) | 50.13 | 78.32 | 2.3040 | 75.50 | training in progress; not a result |
|
| 211 |
+
| 128M final (16Γ768) | 360,000 (47.4%) | 50.84 | 78.01 | 2.3011 | 75.64 | training in progress; not a result |
|
| 212 |
|
| 213 |
Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
|
| 214 |
|
|
|
|
| 228 |
| 128M v1 (`v1_128m/`) | 240,000 | 50.00 | 76.67 | 2.3094 | 74.89 |
|
| 229 |
| 128M v1 (`v1_128m/`) | 280,000 | 51.14 | 77.10 | 2.2950 | 75.46 |
|
| 230 |
| 128M v1 (`v1_128m/`) | 320,000 | 52.06 | 77.24 | 2.2812 | 75.85 |
|
| 231 |
+
| 128M v1 (`v1_128m/`) | 360,000 | 52.53 | 77.12 | 2.2731 | 75.99 |
|
| 232 |
| 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
|
| 233 |
|
| 234 |
+
At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl β0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP β1.72, byte_ppl β0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl β0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP β1.71, byte_ppl β0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy β1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy β0.25, BLiMP β1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. At step 360,000 it is 1.01 eff behind (74.82 vs 75.83: ARC-Easy β1.90, BLiMP β0.89, byte_ppl +0.0171), with v1's learning rate at about 12% of peak versus 59%. At step 240,000 the final 128M run is 0.08 eff behind its v1 (74.81 vs 74.89: ARC-Easy β1.09, BLiMP +1.00, byte_ppl +0.0193), with v1's learning rate at about 41% of peak versus 80%. At step 400,000, where v1 64M ended, the final 64M run is 0.57 eff behind v1's result (75.24 vs 75.81: ARC-Easy β0.97, BLiMP β0.52, byte_ppl +0.0212); v1 was then at its minimum learning rate (10% of peak), while the final run is at about 52% of peak with 360,000 steps still to go. At step 280,000 the final 128M run is +0.17 eff ahead of its v1 (75.63 vs 75.46: ARC-Easy β0.97, BLiMP +1.64, byte_ppl +0.0195), with v1's learning rate at about 29% of peak versus 73%. At step 320,000 it is 0.41 eff behind (75.44 vs 75.85: ARC-Easy β1.60, BLiMP +0.55, byte_ppl +0.0214), with v1's learning rate at about 19% of peak versus 66%. At step 360,000 it is 0.35 eff behind (75.64 vs 75.99: ARC-Easy β1.69, BLiMP +0.89, byte_ppl +0.0280), with v1's learning rate at about 12% of peak versus 59%. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
|
| 235 |
|
| 236 |
**Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
|
| 237 |
|