README for building pair 1 (s12A arcmix pool, s12B arcmix+edu): why they exist, setup
Browse files- s12A_arcmix_pool/README.md +12 -0
- s12B_arcmix_edu/README.md +12 -0
s12A_arcmix_pool/README.md
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# s12A_arcmix_pool — building pair 1, arm A: ARC-MIX sample (2.62B-token pool)
|
| 2 |
+
|
| 3 |
+
**Why it exists:** the reference arm of the first building pair. It continues the 64M flagship at a higher learning rate on a 2.62B-token random sample of ARC-MIX documents (drawn uniformly over documents, without replacement) (the same pool as `r4K_anchor_arcmix_qa1/`), so that both arms draw from pools of the same size and repeat data at the same rate. Arm B (`s12B_arcmix_edu/`) differs only in its data.
|
| 4 |
+
|
| 5 |
+
## Setup
|
| 6 |
+
|
| 7 |
+
- **Base model:** GoLLeM-v5 64M flagship (`v1_muon/ckpt_400k.pt`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Both arms of this pair start from the flagship checkpoint at step 400,000.
|
| 8 |
+
- **Schedule:** warmup-stable-decay. The learning rate is re-warmed over 2,000 steps from 6e-5 to 3e-4, held constant to step 559,980 and decayed to 6e-5 by step 600,000. Batch 32 × 1024 tokens, seed 1337, optimizer state from the checkpoint. The only difference between the two arms is the training data.
|
| 9 |
+
- **Method:** the pair is compared every 20,000 steps on held-out selection sets (ARC-Easy validation, WikiText-2 validation with overlapping articles removed, half of BLiMP). The better arm continues; a tie keeps arm A. Final numbers are reported on the untouched halves and the board test sets.
|
| 10 |
+
- **Status:** research checkpoints in progress, **not a leaderboard submission**. Checkpoints saved during the constant-LR phase are not decayed and are expected to score below decayed models; do not compare them directly.
|
| 11 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
|
| 12 |
+
- **Data:** ARC-MIX, see the root card of this repository.
|
s12B_arcmix_edu/README.md
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# s12B_arcmix_edu — building pair 1, arm B: ARC-MIX + educational web text (2.55B-token pool)
|
| 2 |
+
|
| 3 |
+
**Why it exists:** tests whether adding educational web text helps the 64M flagship at a higher learning rate. The pool is 45% ARC-MIX and 55% FineWeb-Edu text (the same blend as `r3A_arcmix_edu_clean/`), 2.55B tokens. The FineWeb-Edu part is new to the model while ARC-MIX was already seen during pretraining, so a difference between the arms measures "fresh educational data" rather than data quality alone. Arm A (`s12A_arcmix_pool/`) differs only in its data.
|
| 4 |
+
|
| 5 |
+
## Setup
|
| 6 |
+
|
| 7 |
+
- **Base model:** GoLLeM-v5 64M flagship (`v1_muon/ckpt_400k.pt`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Both arms of this pair start from the flagship checkpoint at step 400,000.
|
| 8 |
+
- **Schedule:** warmup-stable-decay. The learning rate is re-warmed over 2,000 steps from 6e-5 to 3e-4, held constant to step 559,980 and decayed to 6e-5 by step 600,000. Batch 32 × 1024 tokens, seed 1337, optimizer state from the checkpoint. The only difference between the two arms is the training data.
|
| 9 |
+
- **Method:** the pair is compared every 20,000 steps on held-out selection sets (ARC-Easy validation, WikiText-2 validation with overlapping articles removed, half of BLiMP). The better arm continues; a tie keeps arm A. Final numbers are reported on the untouched halves and the board test sets.
|
| 10 |
+
- **Status:** research checkpoints in progress, **not a leaderboard submission**. Checkpoints saved during the constant-LR phase are not decayed and are expected to score below decayed models; do not compare them directly.
|
| 11 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
|
| 12 |
+
- **Data:** 45% ARC-MIX + 55% FineWeb-Edu blend (as in `r3A_arcmix_edu_clean/`); see the root card of this repository.
|