|
Download s12B_arcmix_edu/README.md from SlayerLab/gollem-v5-ckpts: direct link, hf CLI and curl.
- Browser
- Download file 1.93 kB
-
https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/d00278b341ff32c9bc047c56981efcde2ee980cc/s12B_arcmix_edu/README.md
- Command line
-
hf download hf://SlayerLab/gollem-v5-ckpts@d00278b341ff32c9bc047c56981efcde2ee980cc/s12B_arcmix_edu/README.md
-
curl -L -o README.md https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/d00278b341ff32c9bc047c56981efcde2ee980cc/s12B_arcmix_edu/README.md
1.93 kB
s12B_arcmix_edu — building pair 1, arm B: ARC-MIX + educational web text (2.55B-token pool)
Why it exists: tests whether adding educational web text helps the 64M flagship at a higher learning rate. The pool is 45% ARC-MIX and 55% FineWeb-Edu text (the same blend as r3A_arcmix_edu_clean/), 2.55B tokens. The FineWeb-Edu part is new to the model while ARC-MIX was already seen during pretraining, so a difference between the arms measures "fresh educational data" rather than data quality alone. Arm A (s12A_arcmix_pool/) differs only in its data.
Setup
- Base model: GoLLeM-v5 64M flagship (
v1_muon/ckpt_400k.pt, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Both arms of this pair start from the flagship checkpoint at step 400,000. - Schedule: warmup-stable-decay. The learning rate is re-warmed over 2,000 steps from 6e-5 to 3e-4, held constant to step 559,980 and decayed to 6e-5 by step 600,000. Batch 32 × 1024 tokens, seed 1337, optimizer state from the checkpoint. The only difference between the two arms is the training data.
- Method: the pair is compared every 20,000 steps on held-out selection sets (ARC-Easy validation, WikiText-2 validation with overlapping articles removed, half of BLiMP). The better arm continues; a tie keeps arm A. Final numbers are reported on the untouched halves and the board test sets.
- Status: research checkpoints in progress, not a leaderboard submission. Checkpoints saved during the constant-LR phase are not decayed and are expected to score below decayed models; do not compare them directly.
- Format: PyTorch checkpoint dict with
model,opt,step,config;train_gpt_ref.pyin the repository root rebuilds the model fromconfig. - Data: 45% ARC-MIX + 55% FineWeb-Edu blend (as in
r3A_arcmix_edu_clean/); see the root card of this repository.