Maggio33's picture
Model card update (BLiMP recomputed after harness fix; data-attribution experiments) + READMEs for 4 experiment folders
1cd1e4d verified
|
Raw History Blame
2.91 kB

r3A_arcmix_edu_clean — arm A of round 3: ARC-MIX + educational text, clean build

Why it exists: a clean repeat of the first educational-data arm. Same proportion as forkB_arcmix_edu/ (45.1% ARC-MIX, 54.9% educational text, 2.55B-token pool), with the build fixed: ARC-MIX sampled uniformly across the whole corpus at document boundaries, and every document separated by the corpus end-of-text token. It was trained in parallel with arm B (r3B_arcmix_qa2x/) from the same checkpoint.

Results

checkpoint step (read from file) ARC-Easy BLiMP WikiText-2 byte-ppl eff (board formula, 62.9M)
ckpt_340k.pt 340000 47.35 76.13 2.3801 75.69
ckpt_360k.pt 360000 47.18 76.01 2.3839 75.58
ckpt_380k.pt 380000 47.10 76.04 2.3787 75.57
ckpt_400k.pt 400000 47.31 76.22 2.3762 75.71

Verdict (A − B, paired bootstrap on eff): +0.46 / −0.13 / −0.02 / +0.11 at 340k / 360k / 380k / 400k; at 400k +0.11 [−0.30, +0.51]. No difference on eff. Across all four checkpoints, arm A has slightly worse WikiText-2 byte-perplexity (a consistent effect of the educational data at this dose) and slightly higher BLiMP (within noise). Conclusion recorded for the study: at this data dose, changing data only in the last 80k steps moves eff by less than about 0.5.

Common setup

  • Base model: GoLLeM-v5 64M flagship (v1_muon/, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs.
  • Method: two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data.
  • Evaluation: glint_parity_eval.py in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
  • Reference: flagship v1_muon/ckpt_400k.pt scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate).
  • Status: research checkpoint, not a leaderboard submission. No arm of this study beats the flagship beyond run-to-run noise.
  • Format: PyTorch checkpoint dict with model, opt, step, config; train_gpt_ref.py in the repository root rebuilds the model from config.