# final_128m_16x768 — GoLLeM-v5 128M final (16 layers × 768) **Why it exists:** a final model of the GoLLeM-v5 series. It is trained from scratch on the ARC-MIX corpus with the recipe of our 64M/128M models and a trainer that draws training windows without accidental repetition; the data composition was chosen by a pre-registered comparison against FineWeb-Edu at 32M (a tie, so the existing corpus was kept). ## Setup - **Model:** 122.8M parameters, 16 layers, d_model 768, 12 heads (head dim 64), RoPE (theta 100,000), SwiGLU (FFN multiplier 2.667), RMSNorm, QK-norm, value residual; BPE tokenizer with 12,288 tokens (`tokenizer.json` in the repository root), context 1024. - **Recipe:** Muon (hidden 2-D weights) + AdamW, peak learning rate 6e-4 (Muon 0.02), 2,000 warmup steps, cosine decay to 6e-5, batch 32 × 1024 tokens, 760,000 steps (24.9B tokens, about 2.65 passes over the pool), seed 1337. Each pass over the pool uses a new permutation of training windows. - **Data:** ARC-MIX, 9.39B tokens after scanning against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching (matching documents were removed); documents containing the marker ‘CC BY-NC-SA’ were also removed. See the root card of this repository for the composition of ARC-MIX. **Attributions:** FineWeb-Edu (HuggingFaceFW/fineweb-edu, ODC-BY 1.0); OpenStax textbooks (CC-BY 4.0, © Rice University, openstax.org; titles and editions listed in OPENSTAX_ATTRIBUTION.md in this folder); minimal-en-corpus-5b (SlayerLab/minimal-en-corpus-5b; see its card for component licences). - **Checkpoints:** every 20,000 steps. **The result is the final checkpoint (step 760,000); intermediate checkpoints are public but are not used for selection or for reporting.** - **Status:** finished 2026-09-28 13:35 UTC. **Result = `ckpt_760k.pt`** (LFS sha256 `95d8b43f…`; the tail-average rule, fixed before the end of training, gave +0.255 < 0.3 on the selection sets): Glint-1.3 `benchmark.py` protocol **ARC-Easy 53.24 / BLiMP 79.09 / WikiText-2 byte_ppl 2.2538 → eff 76.94** (final checkpoint (result)). Against the 128M v1 model (75.84): +1.10 eff at 1.9× the tokens; single seed per run. - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.