Maggio33 commited on
Commit
963440c
·
verified ·
1 Parent(s): acbb407

Card: 128M final result (ckpt_760k, eff 76.94), progress rows 560k-760k, resume wording

Browse files
Files changed (2) hide show
  1. README.md +20 -8
  2. final_128m_16x768/README.md +1 -1
README.md CHANGED
@@ -31,7 +31,7 @@ Every run folder holds plain PyTorch checkpoints (`ckpt.pt` or `ckpt_<N>k.pt`, w
31
  | folder | model | status | result checkpoint | live metrics |
32
  |---|---|---|---|---|
33
  | `final_64m_14x576/` | GoLLeM-v5 64M final, 14 layers / d_model 576 / 9 heads, 62.9M | finished 2026-09-27, checkpoints every 20k steps | `ckpt_760k.pt` (eff 76.07) | [track](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) |
34
- | `final_128m_16x768/` | GoLLeM-v5 128M final, 16 layers / d_model 768 / 12 heads, 122.8M | training, checkpoint every 20k steps | `ckpt_760k.pt` or the tail average (when finished) | [track](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
35
 
36
  **Published models (board entries and references)**
37
 
@@ -138,9 +138,9 @@ Related repositories: the ARC-MIX corpus card [`SlayerLab/gollem-v5-arcmix-9b`](
138
 
139
  ★ 64M flagship (v1 Muon) = Qwen3-style decoder + value residuals, **Muon** optimizer, ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). Recomputed with the fixed harness: **ARC-Easy 47.94 / BLiMP 75.83 / WikiText-2 byte_ppl 2.372 (BPB 1.246) → eff 75.81.** The maintainer's independent re-benchmark (PR #78 and discussion #1; checkpoint sha256 `59f982c1…` matched) reproduced ARC-Easy 47.94 exactly; its BLiMP 75.99 matches our pre-fix 73,000-pair variant to two decimals. Earlier versions of this card cited eff 77.51 (byte_ppl 2.016, from a wrong bytes-per-token factor 4.755 instead of 3.8605) and later eff ~75.9 (BLiMP 75.99); both are superseded. Muon was at least as good as AdamW in a clean 64M A/B (identical data, architecture and seed; only the optimizer differs); those numbers pre-date the BLiMP fix.
140
 
141
- ## Final GoLLeM-v5 runs (64M finished, 128M in progress)
142
 
143
- Two final models were trained with a recipe fixed before the runs started. The 64M run finished on 2026-09-27; the 128M run is still training. Neither the shape nor the data is chosen by looking at intermediate results; the two candidate changes tested beforehand (FineWeb-Edu instead of ARC-MIX, a deeper shape) did not pass their pre-registered rules (see the repository map).
144
 
145
  | | GoLLeM-v5 64M final | GoLLeM-v5 128M final |
146
  |---|---|---|
@@ -151,21 +151,23 @@ Two final models were trained with a recipe fixed before the runs started. The 6
151
  | checkpoints | every 20,000 steps (38 in total) | every 20,000 steps (38 in total) |
152
  | hardware | 1× RTX 5090 (Xeon Gold 6530 host to step 300,000; new pod with Ryzen 9 9950X host after the resume) | 1× RTX 5090 (Ryzen 9 9950X host; new pod after the resume) |
153
  | measured speed | ~202k tokens/s (to step 300,000), ~224k tokens/s after the resume | ~140k tokens/s |
154
- | started / end (UTC) | 2026-09-26 09:09 / 2026-09-27 19:41 (finished) | 2026-09-26 09:09 / ~2026-09-28 13:35 (expected) |
155
  | live metrics | [a7e6791b-30c5-4546-b157-8aac12c4d6fd](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) | [d96091ad-922b-404c-8e8d-6ddaa5fa090a](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
156
 
157
  - **Architecture:** Qwen3-style decoder as in the 64M flagship: RMSNorm, RoPE (θ = 100,000), SwiGLU (ratio 2.667), QK-norm, value residuals, context 1,024, vocabulary 12,288 (BPE).
158
  - **Optimization:** Muon (lr 0.02) for the 2-D hidden weights and AdamW for the rest (betas 0.9 / 0.95, weight decay 0.1); peak learning rate 6e-4, 2,000 warmup steps, cosine decay to 6e-5 at step 760,000; batch 32 × 1,024 tokens; seed 1337; no z-loss, no logit cap.
159
- - **Trainer:** `train_gpt_ref_r6.py` (sha256 `0223c083…`, in the repository root). Each training window is drawn once per pass over the pool; the earlier runs sampled windows with replacement, and a resume could replay early windows. Resume from a checkpoint is exact. A supervisor restarts the trainer from the last checkpoint only after an out-of-memory exit (at most three times).
160
  - **Data (pool sha256 `ecfd0a40…`, 9,391,706,576 tokens):** ARC-MIX, built from the base corpus (5,396,605,407 tokens) + FineWeb-Edu (2,756,248,149) + 58 OpenStax CC BY 4.0 textbooks ×4 (140,742,204) + extra copies of ARC-relevant FineWeb-Edu documents (1,123,440,072) = 9,417,035,832 tokens. From it we removed every document that matched WikiText-2 test/validation or ARC validation/test in a normalized 13-gram and short-question scan (4,362 documents); documents containing the marker ‘CC BY-NC-SA’ were also removed (1,175 documents). 55 OpenStax titles remain; they are listed in `OPENSTAX_ATTRIBUTION.md` in each final folder. The pool was built independently on both machines with identical hashes.
161
  - **Measurement:** the Glint-1.3 `benchmark.py` protocol: BLiMP 67,000 pairs (first token not scored), ARC-Easy test 2,376 questions (bare prompt, LL(question + choice) − LL(question)), WikiText-2 test in 256-token windows, byte perplexity with 3.8605 bytes per token for this tokenizer. eff = mean(BLiMP, ARC-Easy, WikiScore) × size multiplier (1.03646 for the 64M, 1.00839 for the 128M, computed from the declared sizes 62.9M and 122.8M as the board does), with the board's own constants.
162
  - **Resume after the loss of the pods:** the original pods were lost on 2026-09-26 at 22:38 UTC. Both runs were resumed on new pods on 2026-09-27 at 00:58 UTC from the step-300,000 (64M) and step-200,000 (128M) checkpoints on the Hub (SHA-256 checked against the Hub's LFS hash), with the same trainer, flags and pool (rebuilt on the new pods; SHA-256 `ecfd0a40…` checked before the start). Steps 300,000–301,000 (64M) and 200,000–208,000 (128M) were recomputed on the new hardware; the tracker keeps the points it already had for those steps. On those recomputed steps the training loss matches the original run to within 4×10⁻⁵ (5 logged steps for the 64M, 40 for the 128M), so the resumed runs see the same data in the same order as an uninterrupted run.
163
- - **Result = the last checkpoint (step 760,000)**, unless a rule written before the end of training chooses the average of the 720k, 740k and 760k checkpoints: the average is used only if it gains more than 0.3 eff over the last checkpoint on our selection sets (not the board test) and no axis is worse than its noise level. For the 64M run it did not (+0.195), so the 64M result is `ckpt_760k.pt`; for the 128M run the rule is applied after training ends. Other intermediate checkpoints are public but are not used for reporting.
164
  - **Recipe for reproduction:** data mix, training configs and the list of experiments are in [Fabryka-AI-Notes](https://github.com/slayerlabs/Fabryka-AI-Notes/tree/main/papers/slayerlabs/2026-09-26_gollem-v5-report/recipe); `pool_build/` rebuilds the training pool (sha256 `ecfd0a40…`) from the public ARC-MIX file.
165
 
166
  **64M result** (`final_64m_14x576/ckpt_760k.pt`, LFS sha256 `d144c26f…`), Glint-1.3 `benchmark.py` protocol: **ARC-Easy 48.19 / BLiMP 76.16 / WikiText-2 byte_ppl 2.3474 → eff 76.07.** Against the 64M v1 flagship (eff 75.81) this is +0.26, below our 0.6-point decision floor, from a single seed.
167
 
168
- **Progress of the final runs.** Measured with the Glint-1.3 `benchmark.py` protocol (ARC-Easy test 2,376 questions, BLiMP 67,000 pairs, WikiText-2 test) on the checkpoints uploaded so far. Rows marked "not a result" are intermediate checkpoints (for the 128M run: **training in progress; not a result**). The recipe is fixed, so these numbers select nothing; the result is the checkpoint at step 760,000 (or the tail average, see above). The table grows as new checkpoints are measured. The public tracker shows the same board numbers and the training metrics.
 
 
169
 
170
  | run | step (of 760,000) | ARC-Easy | BLiMP | WikiText-2 byte_ppl | eff | status |
171
  |---|---|---|---|---|---|---|
@@ -234,6 +236,17 @@ Two final models were trained with a recipe fixed before the runs started. The 6
234
  | 128M final (16×768) | 500,000 (65.8%) | 51.94 | 78.16 | 2.2758 | 76.13 | training in progress; not a result |
235
  | 128M final (16×768) | 520,000 (68.4%) | 51.64 | 78.93 | 2.2718 | 76.30 | training in progress; not a result |
236
  | 128M final (16×768) | 540,000 (71.1%) | 51.98 | 78.67 | 2.2684 | 76.34 | training in progress; not a result |
 
 
 
 
 
 
 
 
 
 
 
237
 
238
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
239
 
@@ -387,7 +400,6 @@ A generic `lm-eval-harness` run scores BLiMP about 2 pp higher on our 64M (77.84
387
 
388
  ## Roadmap
389
 
390
- - Finish the final 128M run (above) and report its result with the canonical protocol (64M: done, see above).
391
  - Not yet tested, each first as a small from-scratch run with a rule fixed in advance: a larger tokenizer vocabulary, question-answer formatted data from the start of training, and distillation from a larger model.
392
 
393
  ## Limitations
 
31
  | folder | model | status | result checkpoint | live metrics |
32
  |---|---|---|---|---|
33
  | `final_64m_14x576/` | GoLLeM-v5 64M final, 14 layers / d_model 576 / 9 heads, 62.9M | finished 2026-09-27, checkpoints every 20k steps | `ckpt_760k.pt` (eff 76.07) | [track](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) |
34
+ | `final_128m_16x768/` | GoLLeM-v5 128M final, 16 layers / d_model 768 / 12 heads, 122.8M | finished 2026-09-28, checkpoints every 20k steps | `ckpt_760k.pt` (eff 76.94) | [track](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
35
 
36
  **Published models (board entries and references)**
37
 
 
138
 
139
  ★ 64M flagship (v1 Muon) = Qwen3-style decoder + value residuals, **Muon** optimizer, ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). Recomputed with the fixed harness: **ARC-Easy 47.94 / BLiMP 75.83 / WikiText-2 byte_ppl 2.372 (BPB 1.246) → eff 75.81.** The maintainer's independent re-benchmark (PR #78 and discussion #1; checkpoint sha256 `59f982c1…` matched) reproduced ARC-Easy 47.94 exactly; its BLiMP 75.99 matches our pre-fix 73,000-pair variant to two decimals. Earlier versions of this card cited eff 77.51 (byte_ppl 2.016, from a wrong bytes-per-token factor 4.755 instead of 3.8605) and later eff ~75.9 (BLiMP 75.99); both are superseded. Muon was at least as good as AdamW in a clean 64M A/B (identical data, architecture and seed; only the optimizer differs); those numbers pre-date the BLiMP fix.
140
 
141
+ ## Final GoLLeM-v5 runs (64M and 128M finished)
142
 
143
+ Two final models were trained with a recipe fixed before the runs started. The 64M run finished on 2026-09-27 and the 128M run on 2026-09-28. Neither the shape nor the data is chosen by looking at intermediate results; the two candidate changes tested beforehand (FineWeb-Edu instead of ARC-MIX, a deeper shape) did not pass their pre-registered rules (see the repository map).
144
 
145
  | | GoLLeM-v5 64M final | GoLLeM-v5 128M final |
146
  |---|---|---|
 
151
  | checkpoints | every 20,000 steps (38 in total) | every 20,000 steps (38 in total) |
152
  | hardware | 1× RTX 5090 (Xeon Gold 6530 host to step 300,000; new pod with Ryzen 9 9950X host after the resume) | 1× RTX 5090 (Ryzen 9 9950X host; new pod after the resume) |
153
  | measured speed | ~202k tokens/s (to step 300,000), ~224k tokens/s after the resume | ~140k tokens/s |
154
+ | started / end (UTC) | 2026-09-26 09:09 / 2026-09-27 19:41 (finished) | 2026-09-26 09:09 / 2026-09-28 13:35 (finished) |
155
  | live metrics | [a7e6791b-30c5-4546-b157-8aac12c4d6fd](https://track.fabryka.ai/run/a7e6791b-30c5-4546-b157-8aac12c4d6fd) | [d96091ad-922b-404c-8e8d-6ddaa5fa090a](https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a) |
156
 
157
  - **Architecture:** Qwen3-style decoder as in the 64M flagship: RMSNorm, RoPE (θ = 100,000), SwiGLU (ratio 2.667), QK-norm, value residuals, context 1,024, vocabulary 12,288 (BPE).
158
  - **Optimization:** Muon (lr 0.02) for the 2-D hidden weights and AdamW for the rest (betas 0.9 / 0.95, weight decay 0.1); peak learning rate 6e-4, 2,000 warmup steps, cosine decay to 6e-5 at step 760,000; batch 32 × 1,024 tokens; seed 1337; no z-loss, no logit cap.
159
+ - **Trainer:** `train_gpt_ref_r6.py` (sha256 `0223c083…`, in the repository root). Each training window is drawn once per pass over the pool; the earlier runs sampled windows with replacement, and a resume could replay early windows. A resume restores the data position and the optimizer state exactly; on GPU the training losses after a resume match the uninterrupted run to about 1e-4, not bit for bit. A supervisor restarts the trainer from the last checkpoint only after an out-of-memory exit (at most three times).
160
  - **Data (pool sha256 `ecfd0a40…`, 9,391,706,576 tokens):** ARC-MIX, built from the base corpus (5,396,605,407 tokens) + FineWeb-Edu (2,756,248,149) + 58 OpenStax CC BY 4.0 textbooks ×4 (140,742,204) + extra copies of ARC-relevant FineWeb-Edu documents (1,123,440,072) = 9,417,035,832 tokens. From it we removed every document that matched WikiText-2 test/validation or ARC validation/test in a normalized 13-gram and short-question scan (4,362 documents); documents containing the marker ‘CC BY-NC-SA’ were also removed (1,175 documents). 55 OpenStax titles remain; they are listed in `OPENSTAX_ATTRIBUTION.md` in each final folder. The pool was built independently on both machines with identical hashes.
161
  - **Measurement:** the Glint-1.3 `benchmark.py` protocol: BLiMP 67,000 pairs (first token not scored), ARC-Easy test 2,376 questions (bare prompt, LL(question + choice) − LL(question)), WikiText-2 test in 256-token windows, byte perplexity with 3.8605 bytes per token for this tokenizer. eff = mean(BLiMP, ARC-Easy, WikiScore) × size multiplier (1.03646 for the 64M, 1.00839 for the 128M, computed from the declared sizes 62.9M and 122.8M as the board does), with the board's own constants.
162
  - **Resume after the loss of the pods:** the original pods were lost on 2026-09-26 at 22:38 UTC. Both runs were resumed on new pods on 2026-09-27 at 00:58 UTC from the step-300,000 (64M) and step-200,000 (128M) checkpoints on the Hub (SHA-256 checked against the Hub's LFS hash), with the same trainer, flags and pool (rebuilt on the new pods; SHA-256 `ecfd0a40…` checked before the start). Steps 300,000–301,000 (64M) and 200,000–208,000 (128M) were recomputed on the new hardware; the tracker keeps the points it already had for those steps. On those recomputed steps the training loss matches the original run to within 4×10⁻⁵ (5 logged steps for the 64M, 40 for the 128M), so the resumed runs see the same data in the same order as an uninterrupted run.
163
+ - **Result = the last checkpoint (step 760,000)**, unless a rule written before the end of training chooses the average of the 720k, 740k and 760k checkpoints: the average is used only if it gains more than 0.3 eff over the last checkpoint on our selection sets (not the board test) and no axis is worse than its noise level. For the 64M run it did not (+0.195), so the 64M result is `ckpt_760k.pt`; for the 128M run it did not either (+0.255 on the selection sets), so the 128M result is `ckpt_760k.pt`. Other intermediate checkpoints are public but are not used for reporting.
164
  - **Recipe for reproduction:** data mix, training configs and the list of experiments are in [Fabryka-AI-Notes](https://github.com/slayerlabs/Fabryka-AI-Notes/tree/main/papers/slayerlabs/2026-09-26_gollem-v5-report/recipe); `pool_build/` rebuilds the training pool (sha256 `ecfd0a40…`) from the public ARC-MIX file.
165
 
166
  **64M result** (`final_64m_14x576/ckpt_760k.pt`, LFS sha256 `d144c26f…`), Glint-1.3 `benchmark.py` protocol: **ARC-Easy 48.19 / BLiMP 76.16 / WikiText-2 byte_ppl 2.3474 → eff 76.07.** Against the 64M v1 flagship (eff 75.81) this is +0.26, below our 0.6-point decision floor, from a single seed.
167
 
168
+ **128M result** (`final_128m_16x768/ckpt_760k.pt`, LFS sha256 `95d8b43f…`), Glint-1.3 `benchmark.py` protocol: **ARC-Easy 53.24 / BLiMP 79.09 / WikiText-2 byte_ppl 2.2538 → eff 76.94.** Against the 128M v1 model (eff 75.84) this is +1.10 eff at 1.9× the tokens; single seed per run.
169
+
170
+ **Progress of the final runs.** Measured with the Glint-1.3 `benchmark.py` protocol (ARC-Easy test 2,376 questions, BLiMP 67,000 pairs, WikiText-2 test) on the checkpoints uploaded so far. Rows marked "not a result" are intermediate checkpoints (for the 128M run before step 760,000: **training in progress; not a result**). The recipe is fixed, so these numbers select nothing; the result is the checkpoint at step 760,000 (or the tail average, see above). The table grows as new checkpoints are measured. The public tracker shows the same board numbers and the training metrics.
171
 
172
  | run | step (of 760,000) | ARC-Easy | BLiMP | WikiText-2 byte_ppl | eff | status |
173
  |---|---|---|---|---|---|---|
 
236
  | 128M final (16×768) | 500,000 (65.8%) | 51.94 | 78.16 | 2.2758 | 76.13 | training in progress; not a result |
237
  | 128M final (16×768) | 520,000 (68.4%) | 51.64 | 78.93 | 2.2718 | 76.30 | training in progress; not a result |
238
  | 128M final (16×768) | 540,000 (71.1%) | 51.98 | 78.67 | 2.2684 | 76.34 | training in progress; not a result |
239
+ | 128M final (16×768) | 560,000 (73.7%) | 52.44 | 78.79 | 2.2692 | 76.53 | training in progress; not a result |
240
+ | 128M final (16×768) | 580,000 (76.3%) | 52.69 | 79.43 | 2.2628 | 76.84 | training in progress; not a result |
241
+ | 128M final (16×768) | 600,000 (78.9%) | 51.77 | 79.32 | 2.2625 | 76.50 | training in progress; not a result |
242
+ | 128M final (16×768) | 620,000 (81.6%) | 52.57 | 79.31 | 2.2585 | 76.78 | training in progress; not a result |
243
+ | 128M final (16×768) | 640,000 (84.2%) | 53.16 | 79.49 | 2.2578 | 77.04 | training in progress; not a result |
244
+ | 128M final (16×768) | 660,000 (86.8%) | 52.69 | 79.59 | 2.2587 | 76.91 | training in progress; not a result |
245
+ | 128M final (16×768) | 680,000 (89.5%) | 53.28 | 79.28 | 2.2562 | 77.01 | training in progress; not a result |
246
+ | 128M final (16×768) | 700,000 (92.1%) | 53.24 | 78.98 | 2.2563 | 76.90 | training in progress; not a result |
247
+ | 128M final (16×768) | 720,000 (94.7%) | 52.65 | 78.92 | 2.2523 | 76.69 | training in progress; not a result |
248
+ | 128M final (16×768) | 740,000 (97.4%) | 53.20 | 79.48 | 2.2533 | 77.06 | training in progress; not a result |
249
+ | 128M final (16×768) | 760,000 (100.0%) | 53.24 | 79.09 | 2.2538 | 76.94 | final checkpoint (result) |
250
 
251
  Reference, same protocol: the v1 models at the end of their runs (step 400,000) and at the same steps as the final runs.
252
 
 
400
 
401
  ## Roadmap
402
 
 
403
  - Not yet tested, each first as a small from-scratch run with a rule fixed in advance: a larger tokenizer vocabulary, question-answer formatted data from the start of training, and distillation from a larger model.
404
 
405
  ## Limitations
final_128m_16x768/README.md CHANGED
@@ -8,5 +8,5 @@
8
  - **Recipe:** Muon (hidden 2-D weights) + AdamW, peak learning rate 6e-4 (Muon 0.02), 2,000 warmup steps, cosine decay to 6e-5, batch 32 × 1024 tokens, 760,000 steps (24.9B tokens, about 2.65 passes over the pool), seed 1337. Each pass over the pool uses a new permutation of training windows.
9
  - **Data:** ARC-MIX, 9.39B tokens after scanning against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching (matching documents were removed); documents containing the marker ‘CC BY-NC-SA’ were also removed. See the root card of this repository for the composition of ARC-MIX. **Attributions:** FineWeb-Edu (HuggingFaceFW/fineweb-edu, ODC-BY 1.0); OpenStax textbooks (CC-BY 4.0, © Rice University, openstax.org; titles and editions listed in OPENSTAX_ATTRIBUTION.md in this folder); minimal-en-corpus-5b (SlayerLab/minimal-en-corpus-5b; see its card for component licences).
10
  - **Checkpoints:** every 20,000 steps. **The result is the final checkpoint (step 760,000); intermediate checkpoints are public but are not used for selection or for reporting.**
11
- - **Status:** training in progress. The result will be measured with the leaderboard's evaluation protocol on the final checkpoint.
12
  - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
 
8
  - **Recipe:** Muon (hidden 2-D weights) + AdamW, peak learning rate 6e-4 (Muon 0.02), 2,000 warmup steps, cosine decay to 6e-5, batch 32 × 1024 tokens, 760,000 steps (24.9B tokens, about 2.65 passes over the pool), seed 1337. Each pass over the pool uses a new permutation of training windows.
9
  - **Data:** ARC-MIX, 9.39B tokens after scanning against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching (matching documents were removed); documents containing the marker ‘CC BY-NC-SA’ were also removed. See the root card of this repository for the composition of ARC-MIX. **Attributions:** FineWeb-Edu (HuggingFaceFW/fineweb-edu, ODC-BY 1.0); OpenStax textbooks (CC-BY 4.0, © Rice University, openstax.org; titles and editions listed in OPENSTAX_ATTRIBUTION.md in this folder); minimal-en-corpus-5b (SlayerLab/minimal-en-corpus-5b; see its card for component licences).
10
  - **Checkpoints:** every 20,000 steps. **The result is the final checkpoint (step 760,000); intermediate checkpoints are public but are not used for selection or for reporting.**
11
+ - **Status:** finished 2026-09-28 13:35 UTC. **Result = `ckpt_760k.pt`** (LFS sha256 `95d8b43f…`; the tail-average rule, fixed before the end of training, gave +0.255 < 0.3 on the selection sets): Glint-1.3 `benchmark.py` protocol **ARC-Easy 53.24 / BLiMP 79.09 / WikiText-2 byte_ppl 2.2538 → eff 76.94** (final checkpoint (result)). Against the 128M v1 model (75.84): +1.10 eff at 1.9× the tokens; single seed per run.
12
  - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.