|
Download probe_s12B_420k/README.md from SlayerLab/gollem-v5-ckpts: direct link, hf CLI and curl.
- Browser
- Download file 1.11 kB
-
https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/0d7e8417df7dfbc9b28bbeef47d1001678c6f236/probe_s12B_420k/README.md
- Command line
-
hf download hf://SlayerLab/gollem-v5-ckpts@0d7e8417df7dfbc9b28bbeef47d1001678c6f236/probe_s12B_420k/README.md
-
curl -L -o README.md https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/0d7e8417df7dfbc9b28bbeef47d1001678c6f236/probe_s12B_420k/README.md
1.11 kB
probe_s12B_420k — decay probe of s12B_arcmix_edu/ at step 420,000
Why it exists: a short side run that measures what the arm would score if training stopped here. The arm (ARC-MIX + FineWeb-Edu pool, building pair 1 arm B) trains at a constant learning rate, and constant-LR checkpoints are not comparable with decayed models. This probe copies the arm's step-420,000 checkpoint and runs only the decay phase: the learning rate falls linearly from 3e-4 to 6e-5 over 4,000 steps (420,000 → 424,000), on the same data and with the same optimizer state. The arm itself is not affected and keeps training.
- Use: compared with the flagship and with the decay probe of arm A (
probe_s12A_420k/) on the held-out selection sets, it shows how the two arms would compare after decay at this point. - Status: research checkpoint, not a leaderboard submission.
- Format: PyTorch checkpoint dict with
model,opt,step,config;train_gpt_ref.pyin the repository root rebuilds the model fromconfig. - Data: as in
s12B_arcmix_edu/; see the root card of this repository.