# z1_C_arcmix — Z1 arm C: 32M trained from scratch on ARC-MIX **Why it exists:** tests whether training data composition from the first step builds ARC-Easy. Smaller models at the top of the Glint Tiny-ML board were trained from scratch on FineWeb-Edu and score higher on ARC on our evaluation harness than our models trained on ARC-MIX; published small-scale ablations (DataDecide) show the same direction. This arm is the control; it and arm Z (`z1_Z_fineweb_edu/`) use an identical model, recipe and token budget; only the training data differs. ## Setup - **Model:** 31.3M parameters, 15 layers, d_model 384, 6 heads (head dim 64), RoPE, SwiGLU, RMSNorm, QK-norm, value residual; BPE tokenizer with 12,288 tokens (`tokenizer.json` in the repository root). - **Recipe:** the 64M flagship recipe scaled down: Muon (hidden 2-D weights) + AdamW, learning rate 6e-4 with 2,000 warmup steps and cosine decay to 6e-5, batch 32 × 1024 tokens, 150,000 steps (4.9B tokens), seed 1337. Training windows are drawn without replacement (each window at most once), less than one epoch. - **Data:** 5.0B-token uniform sample of ARC-MIX documents (the corpus behind our 32M/64M/128M models; see the root card), scanned against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching; matching documents were removed. - **Checkpoints:** every 30,000 steps (about 1B tokens); the last one is step 150,000. - **Status:** research checkpoints, **not a leaderboard submission**. The comparison is only between arm C and arm Z; the architecture and trainer differ from our published 32M, so their numbers are not directly comparable. - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.