# ctrl_arcmix_resume — control arm (ARC-MIX, unchanged data) **Why it exists:** the reference arm of the first data A/B round. It continues the flagship on the same ARC-MIX corpus the flagship was trained on, so that the other arm (`forkB_arcmix_edu/`) can be compared against "more of the same data". **Data:** ARC-MIX 9.42B-token corpus (the flagship's training data), unchanged. **Important caveat:** when this run was resumed from step 320,000, the trainer restarted its batch sampler from the seed. Because the corpus file is identical to the flagship's, this arm re-drew exactly the same training windows the flagship saw in its first 80k steps. It is therefore a *replay* of early data, not fresh data, and comparisons against it carry that confound. A trainer fix (fast-forwarding the sampler on resume) is prepared for future runs. ## Results | checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) | |---|---|---|---|---|---| | ckpt_360k.pt | 360000 | 46.63 | 75.76 | 2.3727 | 75.33 | | ckpt_400k.pt | 400000 | 47.69 | 76.29 | 2.3717 | 75.88 | ## Common setup - **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs. - **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data. - **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name. - **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate). - **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise. - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.