Maggio33 commited on
Commit
1eadcbc
Β·
verified Β·
1 Parent(s): 693d897

Card: resume after pod loss, 64M 260k-320k, v1 320k reference, new ETA

Browse files
Files changed (1) hide show
  1. README.md +10 -4
README.md CHANGED
@@ -147,9 +147,9 @@ Two final models are training now with a recipe fixed before the runs started. N
147
  | parameters | 62,867,021 | 122,759,951 |
148
  | steps / tokens | 760,000 / 24.9B (β‰ˆ2.65 passes over the pool) | 760,000 / 24.9B (β‰ˆ2.65 passes over the pool) |
149
  | checkpoints | every 20,000 steps (38 in total) | every 20,000 steps (38 in total) |
150
- | hardware | 1Γ— RTX 5090 (Xeon Gold 6530 host) | 1Γ— RTX 5090 (Ryzen 9 9950X host) |
151
- | measured speed | ~202k tokens/s | ~140k tokens/s |
152
- | started / expected end (UTC) | 2026-09-26 09:09 / ~2026-09-27 19:20 | 2026-09-26 09:09 / ~2026-09-28 10:20 |
153
  | live metrics | [a7e6791b-30c5-4546-b157-8aac12c4d6fd](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) | [d96091ad-922b-404c-8e8d-6ddaa5fa090a](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
154
 
155
  - **Architecture:** Qwen3-style decoder as in the 64M flagship: RMSNorm, RoPE (ΞΈ = 100,000), SwiGLU (ratio 2.667), QK-norm, value residuals, context 1,024, vocabulary 12,288 (BPE).
@@ -157,6 +157,7 @@ Two final models are training now with a recipe fixed before the runs started. N
157
  - **Trainer:** each training window is drawn once per pass over the pool; the earlier runs sampled windows with replacement, and a resume could replay early windows. Resume from a checkpoint is exact. A supervisor restarts the trainer from the last checkpoint only after an out-of-memory exit (at most three times).
158
  - **Data (pool sha256 `ecfd0a40…`, 9,391,706,576 tokens):** ARC-MIX, built from the base corpus (5,396,605,407 tokens) + FineWeb-Edu (2,756,248,149) + 58 OpenStax CC BY 4.0 textbooks Γ—4 (140,742,204) + extra copies of ARC-relevant FineWeb-Edu documents (1,123,440,072) = 9,417,035,832 tokens. From it we removed every document that matched WikiText-2 test/validation or ARC validation/test in a normalized 13-gram and short-question scan (4,362 documents); documents containing the marker β€˜CC BY-NC-SA’ were also removed (1,175 documents). 55 OpenStax titles remain; they are listed in `OPENSTAX_ATTRIBUTION.md` in each final folder. The pool was built independently on both machines with identical hashes.
159
  - **Measurement:** the Glint-1.3 `benchmark.py` protocol: BLiMP 67,000 pairs (first token not scored), ARC-Easy test 2,376 questions (bare prompt, LL(question + choice) βˆ’ LL(question)), WikiText-2 test in 256-token windows, byte perplexity with 3.8605 bytes per token for this tokenizer. eff = mean(BLiMP, ARC-Easy, WikiScore) Γ— size multiplier (1.03646 for the 64M, 1.00839 for the 128M, computed from the declared sizes 62.9M and 122.8M as the board does), with the board's own constants.
 
160
  - **Result = the last checkpoint (step 760,000).** Intermediate checkpoints are public but are not used for selection or for reporting.
161
 
162
  **Progress of the final runs.** Measured with the Glint-1.3 `benchmark.py` protocol (ARC-Easy test 2,376 questions, BLiMP 67,000 pairs, WikiText-2 test) on the checkpoints uploaded so far. **Training in progress (step N of 760,000; learning rate still high); not a result.** The recipe is fixed, so these numbers select nothing; the result will be the checkpoint at step 760,000. The table grows as new checkpoints are measured. The tracker's public view shows training metrics (loss, throughput); the evaluation numbers are in this table.
@@ -175,6 +176,10 @@ Two final models are training now with a recipe fixed before the runs started. N
175
  | 64M final (14Γ—576) | 200,000 (26.3%) | 45.50 | 74.72 | 2.4302 | 74.43 | training in progress; not a result |
176
  | 64M final (14Γ—576) | 220,000 (28.9%) | 46.30 | 74.38 | 2.4184 | 74.62 | training in progress; not a result |
177
  | 64M final (14Γ—576) | 240,000 (31.6%) | 46.46 | 74.42 | 2.4185 | 74.69 | training in progress; not a result |
 
 
 
 
178
  | 128M final (16Γ—768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
179
  | 128M final (16Γ—768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
180
  | 128M final (16Γ—768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
@@ -193,6 +198,7 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
193
  | 64M v1 (`v1_muon/`) | 40,000 | 41.41 | 72.61 | 2.5404 | 72.02 |
194
  | 64M v1 (`v1_muon/`) | 80,000 | 43.48 | 75.24 | 2.4922 | 73.76 |
195
  | 64M v1 (`v1_muon/`) | 160,000 | 44.99 | 75.52 | 2.4436 | 74.50 |
 
196
  | 64M v1 (`v1_muon/`), end of run | 400,000 | 47.94 | 75.83 | 2.3718 | 75.81 |
197
  | 128M v1 (`v1_128m/`) | 40,000 | 45.83 | 73.21 | 2.4577 | 71.95 |
198
  | 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.04 |
@@ -201,7 +207,7 @@ Reference, same protocol: the v1 models at the end of their runs (step 400,000)
201
  | 128M v1 (`v1_128m/`) | 200,000 | 49.62 | 76.78 | 2.3344 | 74.74 |
202
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
203
 
204
- At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
205
 
206
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
207
 
 
147
  | parameters | 62,867,021 | 122,759,951 |
148
  | steps / tokens | 760,000 / 24.9B (β‰ˆ2.65 passes over the pool) | 760,000 / 24.9B (β‰ˆ2.65 passes over the pool) |
149
  | checkpoints | every 20,000 steps (38 in total) | every 20,000 steps (38 in total) |
150
+ | hardware | 1Γ— RTX 5090 (Xeon Gold 6530 host to step 300,000; new pod with Ryzen 9 9950X host after the resume) | 1Γ— RTX 5090 (Ryzen 9 9950X host; new pod after the resume) |
151
+ | measured speed | ~202k tokens/s (to step 300,000), ~224k tokens/s after the resume | ~140k tokens/s |
152
+ | started / expected end (UTC) | 2026-09-26 09:09 / ~2026-09-27 19:40 | 2026-09-26 09:09 / ~2026-09-28 13:35 |
153
  | live metrics | [a7e6791b-30c5-4546-b157-8aac12c4d6fd](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) | [d96091ad-922b-404c-8e8d-6ddaa5fa090a](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
154
 
155
  - **Architecture:** Qwen3-style decoder as in the 64M flagship: RMSNorm, RoPE (ΞΈ = 100,000), SwiGLU (ratio 2.667), QK-norm, value residuals, context 1,024, vocabulary 12,288 (BPE).
 
157
  - **Trainer:** each training window is drawn once per pass over the pool; the earlier runs sampled windows with replacement, and a resume could replay early windows. Resume from a checkpoint is exact. A supervisor restarts the trainer from the last checkpoint only after an out-of-memory exit (at most three times).
158
  - **Data (pool sha256 `ecfd0a40…`, 9,391,706,576 tokens):** ARC-MIX, built from the base corpus (5,396,605,407 tokens) + FineWeb-Edu (2,756,248,149) + 58 OpenStax CC BY 4.0 textbooks Γ—4 (140,742,204) + extra copies of ARC-relevant FineWeb-Edu documents (1,123,440,072) = 9,417,035,832 tokens. From it we removed every document that matched WikiText-2 test/validation or ARC validation/test in a normalized 13-gram and short-question scan (4,362 documents); documents containing the marker β€˜CC BY-NC-SA’ were also removed (1,175 documents). 55 OpenStax titles remain; they are listed in `OPENSTAX_ATTRIBUTION.md` in each final folder. The pool was built independently on both machines with identical hashes.
159
  - **Measurement:** the Glint-1.3 `benchmark.py` protocol: BLiMP 67,000 pairs (first token not scored), ARC-Easy test 2,376 questions (bare prompt, LL(question + choice) βˆ’ LL(question)), WikiText-2 test in 256-token windows, byte perplexity with 3.8605 bytes per token for this tokenizer. eff = mean(BLiMP, ARC-Easy, WikiScore) Γ— size multiplier (1.03646 for the 64M, 1.00839 for the 128M, computed from the declared sizes 62.9M and 122.8M as the board does), with the board's own constants.
160
+ - **Resume after the loss of the pods:** the original pods were lost on 2026-09-26 at 22:38 UTC. Both runs were resumed on new pods on 2026-09-27 at 00:58 UTC from the step-300,000 (64M) and step-200,000 (128M) checkpoints on the Hub (SHA-256 checked against the Hub's LFS hash), with the same trainer, flags and pool (rebuilt on the new pods; SHA-256 `ecfd0a40…` checked before the start). Steps 300,000–301,000 (64M) and 200,000–208,000 (128M) were recomputed on the new hardware; the tracker keeps the points it already had for those steps.
161
  - **Result = the last checkpoint (step 760,000).** Intermediate checkpoints are public but are not used for selection or for reporting.
162
 
163
  **Progress of the final runs.** Measured with the Glint-1.3 `benchmark.py` protocol (ARC-Easy test 2,376 questions, BLiMP 67,000 pairs, WikiText-2 test) on the checkpoints uploaded so far. **Training in progress (step N of 760,000; learning rate still high); not a result.** The recipe is fixed, so these numbers select nothing; the result will be the checkpoint at step 760,000. The table grows as new checkpoints are measured. The tracker's public view shows training metrics (loss, throughput); the evaluation numbers are in this table.
 
176
  | 64M final (14Γ—576) | 200,000 (26.3%) | 45.50 | 74.72 | 2.4302 | 74.43 | training in progress; not a result |
177
  | 64M final (14Γ—576) | 220,000 (28.9%) | 46.30 | 74.38 | 2.4184 | 74.62 | training in progress; not a result |
178
  | 64M final (14Γ—576) | 240,000 (31.6%) | 46.46 | 74.42 | 2.4185 | 74.69 | training in progress; not a result |
179
+ | 64M final (14Γ—576) | 260,000 (34.2%) | 46.30 | 74.26 | 2.4077 | 74.61 | training in progress; not a result |
180
+ | 64M final (14Γ—576) | 280,000 (36.8%) | 46.51 | 74.48 | 2.4016 | 74.77 | training in progress; not a result |
181
+ | 64M final (14Γ—576) | 300,000 (39.5%) | 46.59 | 74.41 | 2.4003 | 74.78 | training in progress; not a result |
182
+ | 64M final (14Γ—576) | 320,000 (42.1%), first after the resume | 46.97 | 74.56 | 2.3955 | 74.97 | training in progress; not a result |
183
  | 128M final (16Γ—768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
184
  | 128M final (16Γ—768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
185
  | 128M final (16Γ—768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
 
198
  | 64M v1 (`v1_muon/`) | 40,000 | 41.41 | 72.61 | 2.5404 | 72.02 |
199
  | 64M v1 (`v1_muon/`) | 80,000 | 43.48 | 75.24 | 2.4922 | 73.76 |
200
  | 64M v1 (`v1_muon/`) | 160,000 | 44.99 | 75.52 | 2.4436 | 74.50 |
201
+ | 64M v1 (`v1_muon/`) | 320,000 | 47.22 | 75.89 | 2.3775 | 75.57 |
202
  | 64M v1 (`v1_muon/`), end of run | 400,000 | 47.94 | 75.83 | 2.3718 | 75.81 |
203
  | 128M v1 (`v1_128m/`) | 40,000 | 45.83 | 73.21 | 2.4577 | 71.95 |
204
  | 128M v1 (`v1_128m/`) | 80,000 | 46.38 | 75.53 | 2.4096 | 73.04 |
 
207
  | 128M v1 (`v1_128m/`) | 200,000 | 49.62 | 76.78 | 2.3344 | 74.74 |
208
  | 128M v1 (`v1_128m/`), end of run | 400,000 | 51.94 | 77.26 | 2.2717 | 75.84 |
209
 
210
+ At step 40,000 the final 64M run was +0.99 eff ahead of v1 at the same step (ARC-Easy +2.02, BLiMP +0.85, byte_ppl βˆ’0.0016) and the final 128M run +0.81 eff (ARC-Easy +0.85, BLiMP +1.60, byte_ppl +0.0035). At step 80,000 the 64M runs are level (73.75 vs 73.76: ARC-Easy +1.60, BLiMP βˆ’1.72, byte_ppl βˆ’0.0145), and the final 128M run is +0.79 eff ahead of its v1 (73.83 vs 73.04: ARC-Easy +0.21, BLiMP +2.04, byte_ppl βˆ’0.0137). At step 160,000 the final 64M run is 0.31 eff behind v1 at the same step (74.19 vs 74.50: ARC-Easy +0.80, BLiMP βˆ’1.71, byte_ppl βˆ’0.0020). At step 120,000 the final 128M run is +1.02 eff ahead of its v1 (74.61 vs 73.59: ARC-Easy +0.80, BLiMP +2.32, byte_ppl +0.0132); for v1 128M this is the last checkpoint before its resume. At step 160,000 the final 128M run is +0.95 eff ahead of its v1 (74.81 vs 73.86: ARC-Easy +1.35, BLiMP +1.51, byte_ppl +0.0052). At step 200,000 it is +0.18 eff ahead (74.92 vs 74.74: ARC-Easy βˆ’1.26, BLiMP +1.87, byte_ppl +0.0062). At step 320,000 the final 64M run is 0.60 eff behind v1 at the same step (74.97 vs 75.57: ARC-Easy βˆ’0.25, BLiMP βˆ’1.33, byte_ppl +0.0180); there v1's learning rate was at about 19% of peak versus 66% for the final run. Read this with three caveats: each is a single measurement (one ARC-Easy point is about one standard error at n = 2,376); the v1 pool was not cleaned of the test-set overlaps listed under *Benchmark contamination check*, while the final pool was; and v1 used a shorter cosine schedule (400k steps; at step 160,000 its learning rate was at about 69% of peak versus 91% for the final run), so from roughly 100k steps on the same-step comparison increasingly favours v1.
211
 
212
  **Learning curves** (training in progress; not a result; final 64M to step 240,000, 128M to step 160,000):
213