# s12B2_arcmix_edu — building pair 1, arm B2: ARC-MIX + educational web text, forked from arm A at 420k **Why it exists:** segment 1 of the pair (400k → 420k) ended in a tie, so arm A (`s12A_arcmix_pool/`) stays and the data candidate is re-tested from arm A's checkpoint. This arm is forked from arm A at step 420,000 after the segment-1 tie and trains on the same 2.55B-token pool as `s12B_arcmix_edu/` (45% ARC-MIX, 55% FineWeb-Edu). The FineWeb-Edu part is new to the model while ARC-MIX was already seen during pretraining, so a difference between the arms measures "fresh educational data" rather than data quality alone. Arm A differs only in its data. ## Setup - **Base model:** GoLLeM-v5 64M flagship (`v1_muon/ckpt_400k.pt`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Arm A started from the flagship checkpoint at step 400,000; this arm starts from arm A's checkpoint at step 420,000. - **Schedule:** warmup-stable-decay. The learning rate (re-warmed from 6e-5 to 3e-4 by arm A over steps 400,000-402,000) is held constant at 3e-4 from step 420,000 to 559,980 and decayed to 6e-5 by step 600,000. Batch 32 × 1024 tokens, seed 1337, optimizer state from the checkpoint. The only difference between the two arms is the training data. - **Method:** the pair is compared every 20,000 steps on held-out selection sets (ARC-Easy validation, WikiText-2 validation with overlapping articles removed, half of BLiMP; from segment 2 the ARC axis also includes ARC-Easy train questions not matched by our decontamination scan in either arm's training pool). The decision is made on short decay probes of both arms, not on the constant-LR checkpoints. The better arm continues; a tie keeps arm A. Final numbers are reported on the untouched halves and the board test sets. - **Status:** research checkpoints, **not a leaderboard submission**. Stopped after segment 2 (tie at the decay probes); the last checkpoint is step 460,000. Checkpoints saved during the constant-LR phase are not decayed and are expected to score below decayed models; do not compare them directly. - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`. - **Data:** 45% ARC-MIX + 55% FineWeb-Edu blend (as in `r3A_arcmix_edu_clean/`); see the root card of this repository.