Maggio33's picture
Probe READMEs: decay is 1-sqrt (as in the trainer's WSD), not linear
38bcf33 verified
|
Raw History Blame
1.28 kB

probe_s12A_420k — decay probe of s12A_arcmix_pool/ at step 420,000

Why it exists: a short side run that measures what the arm would score if training stopped here. The arm (ARC-MIX pool, building pair 1 arm A) trains at a constant learning rate, and constant-LR checkpoints are not comparable with decayed models. This probe copies the arm's step-420,000 checkpoint and runs only the decay phase: the learning rate decays from 3e-4 to 6e-5 over 4,000 steps with a 1−sqrt schedule (as in the trainer's WSD decay) (420,000 → 424,000), on the same data and with the same optimizer state. The arm itself is not affected and keeps training.

  • Use: compared with the flagship on the held-out selection sets, it checks that the higher learning rate is not damaging the model (a rule fixed before the result: WikiText-2 validation worse than the flagship by more than 0.010 byte-ppl, or half-BLiMP worse by more than 0.5 pp, means the pair switches to a 2e-4 peak).
  • Status: research checkpoint, not a leaderboard submission.
  • Format: PyTorch checkpoint dict with model, opt, step, config; train_gpt_ref.py in the repository root rebuilds the model from config.
  • Data: as in s12A_arcmix_pool/; see the root card of this repository.