|
Download probe0_s12A_420k/README.md from SlayerLab/gollem-v5-ckpts: direct link, hf CLI and curl.
- Browser
- Download file 1.21 kB
-
https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/eed709d8fff61508e1f88fd08a2dc2f643fc4e58/probe0_s12A_420k/README.md
- Command line
-
hf download hf://SlayerLab/gollem-v5-ckpts@eed709d8fff61508e1f88fd08a2dc2f643fc4e58/probe0_s12A_420k/README.md
-
curl -L -o README.md https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/eed709d8fff61508e1f88fd08a2dc2f643fc4e58/probe0_s12A_420k/README.md
1.21 kB
| # probe0_s12A_420k — decay-to-zero probe of `s12A_arcmix_pool/` at step 420,000 | |
| **Why it exists:** a short side run that measures what the arm would score if training stopped here. The arm (ARC-MIX pool, building pair 1 arm A) trains at a constant learning rate, and constant-LR checkpoints are not comparable with decayed models. This probe copies the arm's step-420,000 checkpoint and runs only the decay phase: the learning rate decays from 3e-4 to 0 over 4,000 steps with a 1−sqrt schedule (as in the trainer's WSD decay) (420,000 → 424,000), on the same data and with the same optimizer state. The arm itself is not affected and keeps training. | |
| - **Use:** identical to `probe_s12A_420k/` (same checkpoint, data, random stream and 4,000 steps) except that the learning rate decays to 0 instead of 6e-5. Comparing the two probes on the held-out selection sets tests whether decaying fully to zero helps. | |
| - **Status:** research checkpoint, **not a leaderboard submission**. | |
| - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`. | |
| - **Data:** as in `s12A_arcmix_pool/`; see the root card of this repository. | |