Maggio33 commited on
Commit
dbcd70b
Β·
verified Β·
1 Parent(s): 5a9cbdf

Card: 128M final @200k vs v1, 64M ETA and speed updated

Browse files
Files changed (1) hide show
  1. README.md +6 -3
README.md CHANGED
@@ -148,8 +148,8 @@ Two final models are training now with a recipe fixed before the runs started. N
148
  | steps / tokens | 760,000 / 24.9B (β‰ˆ2.65 passes over the pool) | 760,000 / 24.9B (β‰ˆ2.65 passes over the pool) |
149
  | checkpoints | every 20,000 steps (38 in total) | every 20,000 steps (38 in total) |
150
  | hardware | 1Γ— RTX 5090 (Xeon Gold 6530 host) | 1Γ— RTX 5090 (Ryzen 9 9950X host) |
151
- | measured speed | ~204k tokens/s | ~140k tokens/s |
152
- | started / expected end (UTC) | 2026-09-26 09:09 / ~2026-09-27 18:45 | 2026-09-26 09:09 / ~2026-09-28 10:20 |
153
  | live metrics | [a7e6791b-30c5-4546-b157-8aac12c4d6fd](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) | [d96091ad-922b-404c-8e8d-6ddaa5fa090a](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
154
 
155
  - **Architecture:** Qwen3-style decoder as in the 64M flagship: RMSNorm, RoPE (ΞΈ = 100,000), SwiGLU (ratio 2.667), QK-norm, value residuals, context 1,024, vocabulary 12,288 (BPE).
@@ -183,6 +183,8 @@ Two final models are training now with a recipe fixed before the runs started. N
183
  | 128M final (16Γ—768) | 120,000 (15.8%) | 48.82 | 77.51 | 2.3760 | 74.61 | training in progress; not a result |
184
  | 128M final (16Γ—768) | 140,000 (18.4%) | 49.71 | 77.58 | 2.3639 | 74.96 | training in progress; not a result |
185
  | 128M final (16Γ—768) | 160,000 (21.1%) | 49.62 | 77.14 | 2.3523 | 74.81 | training in progress; not a result |
 
 
186
 
187
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
188
 
@@ -196,9 +198,10 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
196
  | 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.04 |
197
  | 128M v1 (`v1_128m/`), last checkpoint before its resume | 120,000 | 48.02 | 75.19 | 2.3628 | 73.59 |
198
  | 128M v1 (`v1_128m/`) | 160,000 | 48.27 | 75.63 | 2.3471 | 73.86 |
 
199
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
200
 
201
- At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
202
 
203
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
204
 
 
148
  | steps / tokens | 760,000 / 24.9B (β‰ˆ2.65 passes over the pool) | 760,000 / 24.9B (β‰ˆ2.65 passes over the pool) |
149
  | checkpoints | every 20,000 steps (38 in total) | every 20,000 steps (38 in total) |
150
  | hardware | 1Γ— RTX 5090 (Xeon Gold 6530 host) | 1Γ— RTX 5090 (Ryzen 9 9950X host) |
151
+ | measured speed | ~202k tokens/s | ~140k tokens/s |
152
+ | started / expected end (UTC) | 2026-09-26 09:09 / ~2026-09-27 19:20 | 2026-09-26 09:09 / ~2026-09-28 10:20 |
153
  | live metrics | [a7e6791b-30c5-4546-b157-8aac12c4d6fd](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) | [d96091ad-922b-404c-8e8d-6ddaa5fa090a](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
154
 
155
  - **Architecture:** Qwen3-style decoder as in the 64M flagship: RMSNorm, RoPE (ΞΈ = 100,000), SwiGLU (ratio 2.667), QK-norm, value residuals, context 1,024, vocabulary 12,288 (BPE).
 
183
  | 128M final (16Γ—768) | 120,000 (15.8%) | 48.82 | 77.51 | 2.3760 | 74.61 | training in progress; not a result |
184
  | 128M final (16Γ—768) | 140,000 (18.4%) | 49.71 | 77.58 | 2.3639 | 74.96 | training in progress; not a result |
185
  | 128M final (16Γ—768) | 160,000 (21.1%) | 49.62 | 77.14 | 2.3523 | 74.81 | training in progress; not a result |
186
+ | 128M final (16Γ—768) | 180,000 (23.7%) | 49.41 | 77.65 | 2.3455 | 74.93 | training in progress; not a result |
187
+ | 128M final (16Γ—768) | 200,000 (26.3%) | 48.36 | 78.65 | 2.3406 | 74.92 | training in progress; not a result |
188
 
189
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
190
 
 
198
  | 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.04 |
199
  | 128M v1 (`v1_128m/`), last checkpoint before its resume | 120,000 | 48.02 | 75.19 | 2.3628 | 73.59 |
200
  | 128M v1 (`v1_128m/`) | 160,000 | 48.27 | 75.63 | 2.3471 | 73.86 |
201
+ | 128M v1 (`v1_128m/`) | 200,000 | 49.62 | 76.78 | 2.3344 | 74.74 |
202
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
203
 
204
+ At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
205
 
206
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
207