# v1_128m — GoLLeM-v5 128M (scale reference) **Why it exists:** a scale reference for the 64M flagship (`v1_muon/`). It checks how much of the quality gained by doubling the model survives the size multiplier of the Glint Tiny-ML efficiency score, with the same training data (ARC-MIX) and the same training code family. ## Setup - **Model:** 122.8M parameters, 16 layers, d_model 768, 12 heads (head dim 64), RoPE (theta 100,000), SwiGLU (FFN multiplier 2.667), RMSNorm, QK-norm, value residual; BPE tokenizer with 12,288 tokens (`tokenizer.json` in the repository root), context 1024. - **Training:** Muon (hidden 2-D weights) + AdamW, batch 32 × 1024 tokens, 400,000 steps (13.1B tokens), on ARC-MIX (9.42B tokens), so part of the data is seen a second time. Peak learning rate 6e-4 (Muon 0.02), 2,000 warmup steps, cosine decay to 6e-5. One RTX 5090. The run was resumed at step 120,000 with a trainer version that did not advance the data sampler on resume, so the data windows of steps 0-120k were most likely drawn again after the resume. - **Checkpoints:** every 40,000 steps; the last one is step 400,000. - **Result (step 400,000, our run of the board harness):** ARC-Easy 51.94, BLiMP 77.26, WikiText-2 byte perplexity 2.2717, efficiency score 75.84. The 64M flagship scores 75.81 on the same harness, so at this data and recipe doubling the model adds about as much quality as the size multiplier takes away. - **Status:** research checkpoints, **not a leaderboard submission**. The reported result is the last step of the planned run; intermediate checkpoints are provided for the training trajectory only. - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`. - **Data:** ARC-MIX, see the root card of this repository.