|
Download s12A_arcmix_pool/README.md from SlayerLab/gollem-v5-ckpts: direct link, hf CLI and curl.
- Browser
- Download file 2.11 kB
-
https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/2f7b1bd907e08758945600e2456577c96f2cebd9/s12A_arcmix_pool/README.md
- Command line
-
hf download hf://SlayerLab/gollem-v5-ckpts@2f7b1bd907e08758945600e2456577c96f2cebd9/s12A_arcmix_pool/README.md
-
curl -L -o README.md https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/2f7b1bd907e08758945600e2456577c96f2cebd9/s12A_arcmix_pool/README.md
2.11 kB
| # s12A_arcmix_pool — building pair 1, arm A: ARC-MIX sample (2.62B-token pool) | |
| **Why it exists:** the reference arm of the first building pair. It continues the 64M flagship at a higher learning rate on a 2.62B-token random sample of ARC-MIX documents (drawn uniformly over documents, without replacement) (the same pool as `r4K_anchor_arcmix_qa1/`), so that both arms draw from pools of the same size and repeat data at the same rate. Arm B (`s12B_arcmix_edu/`) differs only in its data. | |
| ## Setup | |
| - **Base model:** GoLLeM-v5 64M flagship (`v1_muon/ckpt_400k.pt`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Both arms of this pair start from the flagship checkpoint at step 400,000. | |
| - **Schedule:** warmup-stable-decay. The learning rate is re-warmed over 2,000 steps from 6e-5 to 3e-4, held constant to step 559,980 and decayed to 6e-5 by step 600,000. Batch 32 × 1024 tokens, seed 1337, optimizer state from the checkpoint. The only difference between the two arms is the training data. | |
| - **Method:** the pair is compared every 20,000 steps on held-out selection sets (ARC-Easy validation, WikiText-2 validation with overlapping articles removed, half of BLiMP; from segment 2 the ARC axis also includes ARC-Easy train questions not matched by our decontamination scan in either arm's training pool). The decision is made on short decay probes of both arms, not on the constant-LR checkpoints. The better arm continues; a tie keeps arm A. Final numbers are reported on the untouched halves and the board test sets. | |
| - **Status:** research checkpoints, **not a leaderboard submission**. Stopped after segment 2 (tie at the decay probes); the last checkpoint is step 460,000. Checkpoints saved during the constant-LR phase are not decayed and are expected to score below decayed models; do not compare them directly. | |
| - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`. | |
| - **Data:** ARC-MIX, see the root card of this repository. | |