Card: 64M final result (760k, eff 76.07), progress to 64M 760k / 128M 540k, tail-averaging rule, r6 trainer, pool_build, recipe link, board status
Browse files- README.md +60 -46
- final_64m_14x576/README.md +4 -3
README.md
CHANGED
|
@@ -24,14 +24,14 @@ The repository holds a controlled scaling study (tokens, width, optimizer) and a
|
|
| 24 |
|
| 25 |
## Repository map (what is where)
|
| 26 |
|
| 27 |
-
Every run folder holds plain PyTorch checkpoints (`ckpt.pt` or `ckpt_<N>k.pt`, where N is the training step in thousands). Folders with a `README.md` (all runs since 2026-09-25) contain the full specification: data pool and its hash, trainer version, seed, steps and tokens. "Result" is the checkpoint that the reported numbers come from; it is
|
| 28 |
|
| 29 |
-
**Final GoLLeM-v5 models
|
| 30 |
|
| 31 |
| folder | model | status | result checkpoint | live metrics |
|
| 32 |
|---|---|---|---|---|
|
| 33 |
-
| `final_64m_14x576/` | GoLLeM-v5 64M final, 14 layers / d_model 576 / 9 heads, 62.9M |
|
| 34 |
-
| `final_128m_16x768/` | GoLLeM-v5 128M final, 16 layers / d_model 768 / 12 heads, 122.8M | training, checkpoint every 20k steps | `ckpt_760k.pt` (when finished) | [track](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
|
| 35 |
|
| 36 |
**Published models (board entries and references)**
|
| 37 |
|
|
@@ -91,6 +91,8 @@ Every run folder holds plain PyTorch checkpoints (`ckpt.pt` or `ckpt_<N>k.pt`, w
|
|
| 91 |
| `eval/` | older evaluation scripts (superseded by `glint_parity_eval.py`) |
|
| 92 |
| `decontam_verify/` | a FineWeb-Edu sample used to check the test-set overlap filter |
|
| 93 |
| `tokenizer.json`, `train_gpt_ref.py`, `glint_parity_eval.py` | tokenizer, model/trainer, board-protocol evaluation |
|
|
|
|
|
|
|
| 94 |
| `glint_*_results.json`, `progress_v5.png`, `scaling_v5.png`, `board_overlay_v5.png` | earlier result files and charts (from before the harness fix; the numbers in this card supersede them) |
|
| 95 |
| other `*.py`, `*.sh` | data-building and pod scripts used for the runs above |
|
| 96 |
|
|
@@ -136,9 +138,9 @@ Related repositories: the ARC-MIX corpus card [`SlayerLab/gollem-v5-arcmix-9b`](
|
|
| 136 |
|
| 137 |
★ 64M flagship (v1 Muon) = Qwen3-style decoder + value residuals, **Muon** optimizer, ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). Recomputed with the fixed harness: **ARC-Easy 47.94 / BLiMP 75.83 / WikiText-2 byte_ppl 2.372 (BPB 1.246) → eff 75.81.** The maintainer's independent re-benchmark (PR #78 and discussion #1; checkpoint sha256 `59f982c1…` matched) reproduced ARC-Easy 47.94 exactly; its BLiMP 75.99 matches our pre-fix 73,000-pair variant to two decimals. Earlier versions of this card cited eff 77.51 (byte_ppl 2.016, from a wrong bytes-per-token factor 4.755 instead of 3.8605) and later eff ~75.9 (BLiMP 75.99); both are superseded. Muon was at least as good as AdamW in a clean 64M A/B (identical data, architecture and seed; only the optimizer differs); those numbers pre-date the BLiMP fix.
|
| 138 |
|
| 139 |
-
## Final GoLLeM-v5 runs (in progress)
|
| 140 |
|
| 141 |
-
Two final models
|
| 142 |
|
| 143 |
| | GoLLeM-v5 64M final | GoLLeM-v5 128M final |
|
| 144 |
|---|---|---|
|
|
@@ -149,55 +151,62 @@ Two final models are training now with a recipe fixed before the runs started. N
|
|
| 149 |
| checkpoints | every 20,000 steps (38 in total) | every 20,000 steps (38 in total) |
|
| 150 |
| hardware | 1× RTX 5090 (Xeon Gold 6530 host to step 300,000; new pod with Ryzen 9 9950X host after the resume) | 1× RTX 5090 (Ryzen 9 9950X host; new pod after the resume) |
|
| 151 |
| measured speed | ~202k tokens/s (to step 300,000), ~224k tokens/s after the resume | ~140k tokens/s |
|
| 152 |
-
| started /
|
| 153 |
| live metrics | [a7e6791b-30c5-4546-b157-8aac12c4d6fd](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) | [d96091ad-922b-404c-8e8d-6ddaa5fa090a](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
|
| 154 |
|
| 155 |
- **Architecture:** Qwen3-style decoder as in the 64M flagship: RMSNorm, RoPE (θ = 100,000), SwiGLU (ratio 2.667), QK-norm, value residuals, context 1,024, vocabulary 12,288 (BPE).
|
| 156 |
- **Optimization:** Muon (lr 0.02) for the 2-D hidden weights and AdamW for the rest (betas 0.9 / 0.95, weight decay 0.1); peak learning rate 6e-4, 2,000 warmup steps, cosine decay to 6e-5 at step 760,000; batch 32 × 1,024 tokens; seed 1337; no z-loss, no logit cap.
|
| 157 |
-
- **Trainer:**
|
| 158 |
- **Data (pool sha256 `ecfd0a40…`, 9,391,706,576 tokens):** ARC-MIX, built from the base corpus (5,396,605,407 tokens) + FineWeb-Edu (2,756,248,149) + 58 OpenStax CC BY 4.0 textbooks ×4 (140,742,204) + extra copies of ARC-relevant FineWeb-Edu documents (1,123,440,072) = 9,417,035,832 tokens. From it we removed every document that matched WikiText-2 test/validation or ARC validation/test in a normalized 13-gram and short-question scan (4,362 documents); documents containing the marker ‘CC BY-NC-SA’ were also removed (1,175 documents). 55 OpenStax titles remain; they are listed in `OPENSTAX_ATTRIBUTION.md` in each final folder. The pool was built independently on both machines with identical hashes.
|
| 159 |
- **Measurement:** the Glint-1.3 `benchmark.py` protocol: BLiMP 67,000 pairs (first token not scored), ARC-Easy test 2,376 questions (bare prompt, LL(question + choice) − LL(question)), WikiText-2 test in 256-token windows, byte perplexity with 3.8605 bytes per token for this tokenizer. eff = mean(BLiMP, ARC-Easy, WikiScore) × size multiplier (1.03646 for the 64M, 1.00839 for the 128M, computed from the declared sizes 62.9M and 122.8M as the board does), with the board's own constants.
|
| 160 |
- **Resume after the loss of the pods:** the original pods were lost on 2026-09-26 at 22:38 UTC. Both runs were resumed on new pods on 2026-09-27 at 00:58 UTC from the step-300,000 (64M) and step-200,000 (128M) checkpoints on the Hub (SHA-256 checked against the Hub's LFS hash), with the same trainer, flags and pool (rebuilt on the new pods; SHA-256 `ecfd0a40…` checked before the start). Steps 300,000–301,000 (64M) and 200,000–208,000 (128M) were recomputed on the new hardware; the tracker keeps the points it already had for those steps. On those recomputed steps the training loss matches the original run to within 4×10⁻⁵ (5 logged steps for the 64M, 40 for the 128M), so the resumed runs see the same data in the same order as an uninterrupted run.
|
| 161 |
-
- **Result = the last checkpoint (step 760,000)
|
|
|
|
| 162 |
|
| 163 |
-
**
|
|
|
|
|
|
|
| 164 |
|
| 165 |
| run | step (of 760,000) | ARC-Easy | BLiMP | WikiText-2 byte_ppl | eff | status |
|
| 166 |
|---|---|---|---|---|---|---|
|
| 167 |
-
| 64M final (14×576) | 20,000 (2.6%) | 40.61 | 72.83 | 2.6070 | 71.66 |
|
| 168 |
-
| 64M final (14×576) | 40,000 (5.3%) | 43.43 | 73.46 | 2.5388 | 73.01 |
|
| 169 |
-
| 64M final (14×576) | 60,000 (7.9%) | 43.77 | 73.41 | 2.4992 | 73.21 |
|
| 170 |
-
| 64M final (14×576) | 80,000 (10.5%) | 45.08 | 73.52 | 2.4777 | 73.75 |
|
| 171 |
-
| 64M final (14×576) | 100,000 (13.2%) | 45.50 | 74.49 | 2.4733 | 74.24 |
|
| 172 |
-
| 64M final (14×576) | 120,000 (15.8%) | 45.58 | 75.15 | 2.4593 | 74.53 |
|
| 173 |
-
| 64M final (14×576) | 140,000 (18.4%) | 44.70 | 74.74 | 2.4501 | 74.11 |
|
| 174 |
-
| 64M final (14×576) | 160,000 (21.1%) | 45.79 | 73.81 | 2.4416 | 74.19 |
|
| 175 |
-
| 64M final (14×576) | 180,000 (23.7%) | 44.99 | 73.72 | 2.4344 | 73.90 |
|
| 176 |
-
| 64M final (14×576) | 200,000 (26.3%) | 45.50 | 74.72 | 2.4302 | 74.43 |
|
| 177 |
-
| 64M final (14×576) | 220,000 (28.9%) | 46.30 | 74.38 | 2.4184 | 74.62 |
|
| 178 |
-
| 64M final (14×576) | 240,000 (31.6%) | 46.46 | 74.42 | 2.4185 | 74.69 |
|
| 179 |
-
| 64M final (14×576) | 260,000 (34.2%) | 46.30 | 74.26 | 2.4077 | 74.61 |
|
| 180 |
-
| 64M final (14×576) | 280,000 (36.8%) | 46.51 | 74.48 | 2.4016 | 74.77 |
|
| 181 |
-
| 64M final (14×576) | 300,000 (39.5%) | 46.59 | 74.41 | 2.4003 | 74.78 |
|
| 182 |
-
| 64M final (14×576) | 320,000 (42.1%), first after the resume | 46.97 | 74.56 | 2.3955 | 74.97 |
|
| 183 |
-
| 64M final (14×576) | 340,000 (44.7%) | 46.80 | 75.05 | 2.3987 | 75.08 |
|
| 184 |
-
| 64M final (14×576) | 360,000 (47.4%) | 45.79 | 75.26 | 2.3917 | 74.82 |
|
| 185 |
-
| 64M final (14×576) | 380,000 (50.0%) | 46.72 | 75.14 | 2.3957 | 75.09 |
|
| 186 |
-
| 64M final (14×576) | 400,000 (52.6%) | 46.97 | 75.31 | 2.3930 | 75.24 |
|
| 187 |
-
| 64M final (14×576) | 420,000 (55.3%) | 47.56 | 75.46 | 2.3865 | 75.51 |
|
| 188 |
-
| 64M final (14×576) | 440,000 (57.9%) | 46.17 | 75.74 | 2.3780 | 75.15 |
|
| 189 |
-
| 64M final (14×576) | 460,000 (60.5%) | 46.30 | 75.47 | 2.3759 | 75.11 |
|
| 190 |
-
| 64M final (14×576) | 480,000 (63.2%) | 46.72 | 75.69 | 2.3734 | 75.33 |
|
| 191 |
-
| 64M final (14×576) | 500,000 (65.8%) | 47.05 | 75.55 | 2.3693 | 75.41 |
|
| 192 |
-
| 64M final (14×576) | 520,000 (68.4%) | 46.80 | 75.60 | 2.3674 | 75.35 |
|
| 193 |
-
| 64M final (14×576) | 540,000 (71.1%) | 46.46 | 75.83 | 2.3678 | 75.31 |
|
| 194 |
-
| 64M final (14×576) | 560,000 (73.7%) | 47.81 | 76.17 | 2.3648 | 75.90 |
|
| 195 |
-
| 64M final (14×576) | 580,000 (76.3%) | 47.47 | 76.38 | 2.3604 | 75.87 |
|
| 196 |
-
| 64M final (14×576) | 600,000 (78.9%) | 47.56 | 76.11 | 2.3561 | 75.81 |
|
| 197 |
-
| 64M final (14×576) | 620,000 (81.6%) | 48.78 | 75.96 | 2.3589 | 76.18 |
|
| 198 |
-
| 64M final (14×576) | 640,000 (84.2%) | 48.15 | 76.64 | 2.3538 | 76.21 |
|
| 199 |
-
| 64M final (14×576) | 660,000 (86.8%) | 47.39 | 76.69 | 2.3557 | 75.96 |
|
| 200 |
-
| 64M final (14×576) | 680,000 (89.5%) | 48.02 | 75.96 | 2.3529 | 75.93 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 201 |
| 128M final (16×768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
|
| 202 |
| 128M final (16×768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
|
| 203 |
| 128M final (16×768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
|
|
@@ -220,6 +229,11 @@ Two final models are training now with a recipe fixed before the runs started. N
|
|
| 220 |
| 128M final (16×768) | 400,000 (52.6%) | 51.73 | 79.38 | 2.2966 | 76.42 | training in progress; not a result |
|
| 221 |
| 128M final (16×768) | 420,000 (55.3%) | 51.73 | 79.26 | 2.2919 | 76.39 | training in progress; not a result |
|
| 222 |
| 128M final (16×768) | 440,000 (57.9%) | 51.30 | 78.63 | 2.2835 | 76.05 | training in progress; not a result |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 223 |
|
| 224 |
Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
|
| 225 |
|
|
@@ -360,7 +374,7 @@ All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e
|
|
| 360 |
|
| 361 |
A generic `lm-eval-harness` run scores BLiMP about 2 pp higher on our 64M (77.84 vs 75.83); the numbers here are the board-comparable ones.
|
| 362 |
|
| 363 |
-
**Board status.** The Glint board still lists the earlier values for our entries (64M: BLiMP 77.84 / byte_ppl 2.016; 32M: BLiMP 73.77; 16M: BLiMP 70.53). The correct values are the ones in this card. With the board's own
|
| 364 |
|
| 365 |
## Key findings
|
| 366 |
|
|
@@ -373,7 +387,7 @@ A generic `lm-eval-harness` run scores BLiMP about 2 pp higher on our 64M (77.84
|
|
| 373 |
|
| 374 |
## Roadmap
|
| 375 |
|
| 376 |
-
- Finish the final
|
| 377 |
- Not yet tested, each first as a small from-scratch run with a rule fixed in advance: a larger tokenizer vocabulary, question-answer formatted data from the start of training, and distillation from a larger model.
|
| 378 |
|
| 379 |
## Limitations
|
|
|
|
| 24 |
|
| 25 |
## Repository map (what is where)
|
| 26 |
|
| 27 |
+
Every run folder holds plain PyTorch checkpoints (`ckpt.pt` or `ckpt_<N>k.pt`, where N is the training step in thousands). Folders with a `README.md` (all runs since 2026-09-25) contain the full specification: data pool and its hash, trainer version, seed, steps and tokens. "Result" is the checkpoint that the reported numbers come from; it is the planned last step (for the final runs, the average of the last three checkpoints if a rule written before the end chooses it), never a selected "best" checkpoint.
|
| 28 |
|
| 29 |
+
**Final GoLLeM-v5 models**
|
| 30 |
|
| 31 |
| folder | model | status | result checkpoint | live metrics |
|
| 32 |
|---|---|---|---|---|
|
| 33 |
+
| `final_64m_14x576/` | GoLLeM-v5 64M final, 14 layers / d_model 576 / 9 heads, 62.9M | finished 2026-09-27, checkpoints every 20k steps | `ckpt_760k.pt` (eff 76.07) | [track](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) |
|
| 34 |
+
| `final_128m_16x768/` | GoLLeM-v5 128M final, 16 layers / d_model 768 / 12 heads, 122.8M | training, checkpoint every 20k steps | `ckpt_760k.pt` or the tail average (when finished) | [track](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
|
| 35 |
|
| 36 |
**Published models (board entries and references)**
|
| 37 |
|
|
|
|
| 91 |
| `eval/` | older evaluation scripts (superseded by `glint_parity_eval.py`) |
|
| 92 |
| `decontam_verify/` | a FineWeb-Edu sample used to check the test-set overlap filter |
|
| 93 |
| `tokenizer.json`, `train_gpt_ref.py`, `glint_parity_eval.py` | tokenizer, model/trainer, board-protocol evaluation |
|
| 94 |
+
| `train_gpt_ref_r6.py` | the trainer of the final runs (sha256 `0223c083…`), published exactly as it ran; code comments are in Polish |
|
| 95 |
+
| `pool_build/` | rebuilds the final training pool bit for bit from ARC-MIX: removed document indices and `rebuild_final_pool.py` |
|
| 96 |
| `glint_*_results.json`, `progress_v5.png`, `scaling_v5.png`, `board_overlay_v5.png` | earlier result files and charts (from before the harness fix; the numbers in this card supersede them) |
|
| 97 |
| other `*.py`, `*.sh` | data-building and pod scripts used for the runs above |
|
| 98 |
|
|
|
|
| 138 |
|
| 139 |
★ 64M flagship (v1 Muon) = Qwen3-style decoder + value residuals, **Muon** optimizer, ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). Recomputed with the fixed harness: **ARC-Easy 47.94 / BLiMP 75.83 / WikiText-2 byte_ppl 2.372 (BPB 1.246) → eff 75.81.** The maintainer's independent re-benchmark (PR #78 and discussion #1; checkpoint sha256 `59f982c1…` matched) reproduced ARC-Easy 47.94 exactly; its BLiMP 75.99 matches our pre-fix 73,000-pair variant to two decimals. Earlier versions of this card cited eff 77.51 (byte_ppl 2.016, from a wrong bytes-per-token factor 4.755 instead of 3.8605) and later eff ~75.9 (BLiMP 75.99); both are superseded. Muon was at least as good as AdamW in a clean 64M A/B (identical data, architecture and seed; only the optimizer differs); those numbers pre-date the BLiMP fix.
|
| 140 |
|
| 141 |
+
## Final GoLLeM-v5 runs (64M finished, 128M in progress)
|
| 142 |
|
| 143 |
+
Two final models were trained with a recipe fixed before the runs started. The 64M run finished on 2026-09-27; the 128M run is still training. Neither the shape nor the data is chosen by looking at intermediate results; the two candidate changes tested beforehand (FineWeb-Edu instead of ARC-MIX, a deeper shape) did not pass their pre-registered rules (see the repository map).
|
| 144 |
|
| 145 |
| | GoLLeM-v5 64M final | GoLLeM-v5 128M final |
|
| 146 |
|---|---|---|
|
|
|
|
| 151 |
| checkpoints | every 20,000 steps (38 in total) | every 20,000 steps (38 in total) |
|
| 152 |
| hardware | 1× RTX 5090 (Xeon Gold 6530 host to step 300,000; new pod with Ryzen 9 9950X host after the resume) | 1× RTX 5090 (Ryzen 9 9950X host; new pod after the resume) |
|
| 153 |
| measured speed | ~202k tokens/s (to step 300,000), ~224k tokens/s after the resume | ~140k tokens/s |
|
| 154 |
+
| started / end (UTC) | 2026-09-26 09:09 / 2026-09-27 19:41 (finished) | 2026-09-26 09:09 / ~2026-09-28 13:35 (expected) |
|
| 155 |
| live metrics | [a7e6791b-30c5-4546-b157-8aac12c4d6fd](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) | [d96091ad-922b-404c-8e8d-6ddaa5fa090a](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
|
| 156 |
|
| 157 |
- **Architecture:** Qwen3-style decoder as in the 64M flagship: RMSNorm, RoPE (θ = 100,000), SwiGLU (ratio 2.667), QK-norm, value residuals, context 1,024, vocabulary 12,288 (BPE).
|
| 158 |
- **Optimization:** Muon (lr 0.02) for the 2-D hidden weights and AdamW for the rest (betas 0.9 / 0.95, weight decay 0.1); peak learning rate 6e-4, 2,000 warmup steps, cosine decay to 6e-5 at step 760,000; batch 32 × 1,024 tokens; seed 1337; no z-loss, no logit cap.
|
| 159 |
+
- **Trainer:** `train_gpt_ref_r6.py` (sha256 `0223c083…`, in the repository root). Each training window is drawn once per pass over the pool; the earlier runs sampled windows with replacement, and a resume could replay early windows. Resume from a checkpoint is exact. A supervisor restarts the trainer from the last checkpoint only after an out-of-memory exit (at most three times).
|
| 160 |
- **Data (pool sha256 `ecfd0a40…`, 9,391,706,576 tokens):** ARC-MIX, built from the base corpus (5,396,605,407 tokens) + FineWeb-Edu (2,756,248,149) + 58 OpenStax CC BY 4.0 textbooks ×4 (140,742,204) + extra copies of ARC-relevant FineWeb-Edu documents (1,123,440,072) = 9,417,035,832 tokens. From it we removed every document that matched WikiText-2 test/validation or ARC validation/test in a normalized 13-gram and short-question scan (4,362 documents); documents containing the marker ‘CC BY-NC-SA’ were also removed (1,175 documents). 55 OpenStax titles remain; they are listed in `OPENSTAX_ATTRIBUTION.md` in each final folder. The pool was built independently on both machines with identical hashes.
|
| 161 |
- **Measurement:** the Glint-1.3 `benchmark.py` protocol: BLiMP 67,000 pairs (first token not scored), ARC-Easy test 2,376 questions (bare prompt, LL(question + choice) − LL(question)), WikiText-2 test in 256-token windows, byte perplexity with 3.8605 bytes per token for this tokenizer. eff = mean(BLiMP, ARC-Easy, WikiScore) × size multiplier (1.03646 for the 64M, 1.00839 for the 128M, computed from the declared sizes 62.9M and 122.8M as the board does), with the board's own constants.
|
| 162 |
- **Resume after the loss of the pods:** the original pods were lost on 2026-09-26 at 22:38 UTC. Both runs were resumed on new pods on 2026-09-27 at 00:58 UTC from the step-300,000 (64M) and step-200,000 (128M) checkpoints on the Hub (SHA-256 checked against the Hub's LFS hash), with the same trainer, flags and pool (rebuilt on the new pods; SHA-256 `ecfd0a40…` checked before the start). Steps 300,000–301,000 (64M) and 200,000–208,000 (128M) were recomputed on the new hardware; the tracker keeps the points it already had for those steps. On those recomputed steps the training loss matches the original run to within 4×10⁻⁵ (5 logged steps for the 64M, 40 for the 128M), so the resumed runs see the same data in the same order as an uninterrupted run.
|
| 163 |
+
- **Result = the last checkpoint (step 760,000)**, unless a rule written before the end of training chooses the average of the 720k, 740k and 760k checkpoints: the average is used only if it gains more than 0.3 eff over the last checkpoint on our selection sets (not the board test) and no axis is worse than its noise level. For the 64M run it did not (+0.195), so the 64M result is `ckpt_760k.pt`; for the 128M run the rule is applied after training ends. Other intermediate checkpoints are public but are not used for reporting.
|
| 164 |
+
- **Recipe for reproduction:** data mix, training configs and the list of experiments are in [Fabryka-AI-Notes](https://github.com/slayerlabs/Fabryka-AI-Notes/tree/main/papers/slayerlabs/2026-09-26_gollem-v5-report/recipe); `pool_build/` rebuilds the training pool (sha256 `ecfd0a40…`) from the public ARC-MIX file.
|
| 165 |
|
| 166 |
+
**64M result** (`final_64m_14x576/ckpt_760k.pt`, LFS sha256 `d144c26f…`), Glint-1.3 `benchmark.py` protocol: **ARC-Easy 48.19 / BLiMP 76.16 / WikiText-2 byte_ppl 2.3474 → eff 76.07.** Against the 64M v1 flagship (eff 75.81) this is +0.26, below our 0.6-point decision floor, from a single seed.
|
| 167 |
+
|
| 168 |
+
**Progress of the final runs.** Measured with the Glint-1.3 `benchmark.py` protocol (ARC-Easy test 2,376 questions, BLiMP 67,000 pairs, WikiText-2 test) on the checkpoints uploaded so far. Rows marked "not a result" are intermediate checkpoints (for the 128M run: **training in progress; not a result**). The recipe is fixed, so these numbers select nothing; the result is the checkpoint at step 760,000 (or the tail average, see above). The table grows as new checkpoints are measured. The public tracker shows the same board numbers and the training metrics.
|
| 169 |
|
| 170 |
| run | step (of 760,000) | ARC-Easy | BLiMP | WikiText-2 byte_ppl | eff | status |
|
| 171 |
|---|---|---|---|---|---|---|
|
| 172 |
+
| 64M final (14×576) | 20,000 (2.6%) | 40.61 | 72.83 | 2.6070 | 71.66 | intermediate checkpoint; not a result |
|
| 173 |
+
| 64M final (14×576) | 40,000 (5.3%) | 43.43 | 73.46 | 2.5388 | 73.01 | intermediate checkpoint; not a result |
|
| 174 |
+
| 64M final (14×576) | 60,000 (7.9%) | 43.77 | 73.41 | 2.4992 | 73.21 | intermediate checkpoint; not a result |
|
| 175 |
+
| 64M final (14×576) | 80,000 (10.5%) | 45.08 | 73.52 | 2.4777 | 73.75 | intermediate checkpoint; not a result |
|
| 176 |
+
| 64M final (14×576) | 100,000 (13.2%) | 45.50 | 74.49 | 2.4733 | 74.24 | intermediate checkpoint; not a result |
|
| 177 |
+
| 64M final (14×576) | 120,000 (15.8%) | 45.58 | 75.15 | 2.4593 | 74.53 | intermediate checkpoint; not a result |
|
| 178 |
+
| 64M final (14×576) | 140,000 (18.4%) | 44.70 | 74.74 | 2.4501 | 74.11 | intermediate checkpoint; not a result |
|
| 179 |
+
| 64M final (14×576) | 160,000 (21.1%) | 45.79 | 73.81 | 2.4416 | 74.19 | intermediate checkpoint; not a result |
|
| 180 |
+
| 64M final (14×576) | 180,000 (23.7%) | 44.99 | 73.72 | 2.4344 | 73.90 | intermediate checkpoint; not a result |
|
| 181 |
+
| 64M final (14×576) | 200,000 (26.3%) | 45.50 | 74.72 | 2.4302 | 74.43 | intermediate checkpoint; not a result |
|
| 182 |
+
| 64M final (14×576) | 220,000 (28.9%) | 46.30 | 74.38 | 2.4184 | 74.62 | intermediate checkpoint; not a result |
|
| 183 |
+
| 64M final (14×576) | 240,000 (31.6%) | 46.46 | 74.42 | 2.4185 | 74.69 | intermediate checkpoint; not a result |
|
| 184 |
+
| 64M final (14×576) | 260,000 (34.2%) | 46.30 | 74.26 | 2.4077 | 74.61 | intermediate checkpoint; not a result |
|
| 185 |
+
| 64M final (14×576) | 280,000 (36.8%) | 46.51 | 74.48 | 2.4016 | 74.77 | intermediate checkpoint; not a result |
|
| 186 |
+
| 64M final (14×576) | 300,000 (39.5%) | 46.59 | 74.41 | 2.4003 | 74.78 | intermediate checkpoint; not a result |
|
| 187 |
+
| 64M final (14×576) | 320,000 (42.1%), first after the resume | 46.97 | 74.56 | 2.3955 | 74.97 | intermediate checkpoint; not a result |
|
| 188 |
+
| 64M final (14×576) | 340,000 (44.7%) | 46.80 | 75.05 | 2.3987 | 75.08 | intermediate checkpoint; not a result |
|
| 189 |
+
| 64M final (14×576) | 360,000 (47.4%) | 45.79 | 75.26 | 2.3917 | 74.82 | intermediate checkpoint; not a result |
|
| 190 |
+
| 64M final (14×576) | 380,000 (50.0%) | 46.72 | 75.14 | 2.3957 | 75.09 | intermediate checkpoint; not a result |
|
| 191 |
+
| 64M final (14×576) | 400,000 (52.6%) | 46.97 | 75.31 | 2.3930 | 75.24 | intermediate checkpoint; not a result |
|
| 192 |
+
| 64M final (14×576) | 420,000 (55.3%) | 47.56 | 75.46 | 2.3865 | 75.51 | intermediate checkpoint; not a result |
|
| 193 |
+
| 64M final (14×576) | 440,000 (57.9%) | 46.17 | 75.74 | 2.3780 | 75.15 | intermediate checkpoint; not a result |
|
| 194 |
+
| 64M final (14×576) | 460,000 (60.5%) | 46.30 | 75.47 | 2.3759 | 75.11 | intermediate checkpoint; not a result |
|
| 195 |
+
| 64M final (14×576) | 480,000 (63.2%) | 46.72 | 75.69 | 2.3734 | 75.33 | intermediate checkpoint; not a result |
|
| 196 |
+
| 64M final (14×576) | 500,000 (65.8%) | 47.05 | 75.55 | 2.3693 | 75.41 | intermediate checkpoint; not a result |
|
| 197 |
+
| 64M final (14×576) | 520,000 (68.4%) | 46.80 | 75.60 | 2.3674 | 75.35 | intermediate checkpoint; not a result |
|
| 198 |
+
| 64M final (14×576) | 540,000 (71.1%) | 46.46 | 75.83 | 2.3678 | 75.31 | intermediate checkpoint; not a result |
|
| 199 |
+
| 64M final (14×576) | 560,000 (73.7%) | 47.81 | 76.17 | 2.3648 | 75.90 | intermediate checkpoint; not a result |
|
| 200 |
+
| 64M final (14×576) | 580,000 (76.3%) | 47.47 | 76.38 | 2.3604 | 75.87 | intermediate checkpoint; not a result |
|
| 201 |
+
| 64M final (14×576) | 600,000 (78.9%) | 47.56 | 76.11 | 2.3561 | 75.81 | intermediate checkpoint; not a result |
|
| 202 |
+
| 64M final (14×576) | 620,000 (81.6%) | 48.78 | 75.96 | 2.3589 | 76.18 | intermediate checkpoint; not a result |
|
| 203 |
+
| 64M final (14×576) | 640,000 (84.2%) | 48.15 | 76.64 | 2.3538 | 76.21 | intermediate checkpoint; not a result |
|
| 204 |
+
| 64M final (14×576) | 660,000 (86.8%) | 47.39 | 76.69 | 2.3557 | 75.96 | intermediate checkpoint; not a result |
|
| 205 |
+
| 64M final (14×576) | 680,000 (89.5%) | 48.02 | 75.96 | 2.3529 | 75.93 | intermediate checkpoint; not a result |
|
| 206 |
+
| 64M final (14×576) | 700,000 (92.1%) | 48.19 | 76.07 | 2.3513 | 76.03 | intermediate checkpoint; not a result |
|
| 207 |
+
| 64M final (14×576) | 720,000 (94.7%) | 48.61 | 75.93 | 2.3489 | 76.13 | intermediate checkpoint; not a result |
|
| 208 |
+
| 64M final (14×576) | 740,000 (97.4%) | 47.98 | 76.16 | 2.3475 | 76.00 | intermediate checkpoint; not a result |
|
| 209 |
+
| 64M final (14×576) | 760,000 (100.0%) | 48.19 | 76.16 | 2.3474 | 76.07 | final checkpoint (result) |
|
| 210 |
| 128M final (16×768) | 20,000 (2.6%) | 43.90 | 73.91 | 2.5383 | 71.34 | training in progress; not a result |
|
| 211 |
| 128M final (16×768) | 40,000 (5.3%) | 46.68 | 74.81 | 2.4612 | 72.77 | training in progress; not a result |
|
| 212 |
| 128M final (16×768) | 60,000 (7.9%) | 47.43 | 75.24 | 2.4151 | 73.28 | training in progress; not a result |
|
|
|
|
| 229 |
| 128M final (16×768) | 400,000 (52.6%) | 51.73 | 79.38 | 2.2966 | 76.42 | training in progress; not a result |
|
| 230 |
| 128M final (16×768) | 420,000 (55.3%) | 51.73 | 79.26 | 2.2919 | 76.39 | training in progress; not a result |
|
| 231 |
| 128M final (16×768) | 440,000 (57.9%) | 51.30 | 78.63 | 2.2835 | 76.05 | training in progress; not a result |
|
| 232 |
+
| 128M final (16×768) | 460,000 (60.5%) | 52.15 | 78.60 | 2.2793 | 76.34 | training in progress; not a result |
|
| 233 |
+
| 128M final (16×768) | 480,000 (63.2%) | 50.13 | 78.66 | 2.2777 | 75.69 | training in progress; not a result |
|
| 234 |
+
| 128M final (16×768) | 500,000 (65.8%) | 51.94 | 78.16 | 2.2758 | 76.13 | training in progress; not a result |
|
| 235 |
+
| 128M final (16×768) | 520,000 (68.4%) | 51.64 | 78.93 | 2.2718 | 76.30 | training in progress; not a result |
|
| 236 |
+
| 128M final (16×768) | 540,000 (71.1%) | 51.98 | 78.67 | 2.2684 | 76.34 | training in progress; not a result |
|
| 237 |
|
| 238 |
Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
|
| 239 |
|
|
|
|
| 374 |
|
| 375 |
A generic `lm-eval-harness` run scores BLiMP about 2 pp higher on our 64M (77.84 vs 75.83); the numbers here are the board-comparable ones.
|
| 376 |
|
| 377 |
+
**Board status.** The Glint board still lists the earlier values for our entries (64M: BLiMP 77.84 / byte_ppl 2.016; 32M: BLiMP 73.77; 16M: BLiMP 70.53). The correct values are the ones in this card; a correction was submitted on 2026-09-27 ([discussion #81](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard/discussions/81)) that fixes the 32M and 16M BLiMP values and replaces the 64M entry with the final 64M model, and the board shows the old values until it is merged. With the board's own code (`eff = mean(BLiMP, ARC-Easy, normalized-WikiText-2) × size multiplier`; board revision 35c5b6f7, 2026-09-27) they give eff **75.80 (64M v1, the entry before #81), 75.40 (32M) and 74.30 (16M)**; the final 64M model gives **76.06** (76.07 in this card's fixed WikiText-2 scale).
|
| 378 |
|
| 379 |
## Key findings
|
| 380 |
|
|
|
|
| 387 |
|
| 388 |
## Roadmap
|
| 389 |
|
| 390 |
+
- Finish the final 128M run (above) and report its result with the canonical protocol (64M: done, see above).
|
| 391 |
- Not yet tested, each first as a small from-scratch run with a rule fixed in advance: a larger tokenizer vocabulary, question-answer formatted data from the start of training, and distillation from a larger model.
|
| 392 |
|
| 393 |
## Limitations
|
final_64m_14x576/README.md
CHANGED
|
@@ -7,6 +7,7 @@
|
|
| 7 |
- **Model:** 62.9M parameters, 14 layers, d_model 576, 9 heads (head dim 64), RoPE (theta 100,000), SwiGLU (FFN multiplier 2.667), RMSNorm, QK-norm, value residual; BPE tokenizer with 12,288 tokens (`tokenizer.json` in the repository root), context 1024.
|
| 8 |
- **Recipe:** Muon (hidden 2-D weights) + AdamW, peak learning rate 6e-4 (Muon 0.02), 2,000 warmup steps, cosine decay to 6e-5, batch 32 × 1024 tokens, 760,000 steps (24.9B tokens, about 2.65 passes over the pool), seed 1337. Each pass over the pool uses a new permutation of training windows.
|
| 9 |
- **Data:** ARC-MIX, 9.39B tokens after scanning against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching (matching documents were removed); documents containing the marker ‘CC BY-NC-SA’ were also removed. See the root card of this repository for the composition of ARC-MIX. **Attributions:** FineWeb-Edu (HuggingFaceFW/fineweb-edu, ODC-BY 1.0); OpenStax textbooks (CC-BY 4.0, © Rice University, openstax.org; titles and editions listed in OPENSTAX_ATTRIBUTION.md in this folder); minimal-en-corpus-5b (SlayerLab/minimal-en-corpus-5b; see its card for component licences).
|
| 10 |
-
- **Checkpoints:** every 20,000 steps. **The result is the final checkpoint (step 760,000); intermediate checkpoints are public but are not used for
|
| 11 |
-
- **Status:**
|
| 12 |
-
- **
|
|
|
|
|
|
| 7 |
- **Model:** 62.9M parameters, 14 layers, d_model 576, 9 heads (head dim 64), RoPE (theta 100,000), SwiGLU (FFN multiplier 2.667), RMSNorm, QK-norm, value residual; BPE tokenizer with 12,288 tokens (`tokenizer.json` in the repository root), context 1024.
|
| 8 |
- **Recipe:** Muon (hidden 2-D weights) + AdamW, peak learning rate 6e-4 (Muon 0.02), 2,000 warmup steps, cosine decay to 6e-5, batch 32 × 1024 tokens, 760,000 steps (24.9B tokens, about 2.65 passes over the pool), seed 1337. Each pass over the pool uses a new permutation of training windows.
|
| 9 |
- **Data:** ARC-MIX, 9.39B tokens after scanning against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching (matching documents were removed); documents containing the marker ‘CC BY-NC-SA’ were also removed. See the root card of this repository for the composition of ARC-MIX. **Attributions:** FineWeb-Edu (HuggingFaceFW/fineweb-edu, ODC-BY 1.0); OpenStax textbooks (CC-BY 4.0, © Rice University, openstax.org; titles and editions listed in OPENSTAX_ATTRIBUTION.md in this folder); minimal-en-corpus-5b (SlayerLab/minimal-en-corpus-5b; see its card for component licences).
|
| 10 |
+
- **Checkpoints:** every 20,000 steps. **The result is the final checkpoint `ckpt_760k.pt` (step 760,000).** A rule written before the end of training would have used the average of the 720k, 740k and 760k checkpoints only if it gained more than 0.3 eff on our selection sets (not the board test) with no axis worse than its noise level; it gained +0.195, so the single last checkpoint is the result. Other intermediate checkpoints are public but are not used for reporting.
|
| 11 |
+
- **Status:** finished on 2026-09-27 at 19:41 UTC.
|
| 12 |
+
- **Result** (`ckpt_760k.pt`, LFS sha256 `d144c26f…`), Glint-1.3 `benchmark.py` protocol: **ARC-Easy 48.19 / BLiMP 76.16 / WikiText-2 byte_ppl 2.3474 → eff 76.07.** Against the 64M v1 flagship (eff 75.81) this is +0.26, below our 0.6-point decision floor, from a single seed.
|
| 13 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`; `train_gpt_ref_r6.py` (sha256 `0223c083…`) is the trainer that produced this run.
|