README for round-4 arms (r4B seed control, r4K anchor): why they exist, results 340-400k
Browse files
r4B_arcmix_qa2x_seed1338/README.md
ADDED
|
@@ -0,0 +1,24 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# r4B_arcmix_qa2x_seed1338 β round 4: arm B of round 3 repeated with a different seed
|
| 2 |
+
|
| 3 |
+
**Why it exists:** measures run-to-run noise. Identical to `r3B_arcmix_qa2x/` (same data, same checkpoint, same recipe) except the random seed (1338 instead of 1337). The difference between the two runs is therefore noise, and it sets the threshold every data comparison in this study has to beat.
|
| 4 |
+
|
| 5 |
+
## Results
|
| 6 |
+
|
| 7 |
+
| checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
|
| 8 |
+
|---|---|---|---|---|---|
|
| 9 |
+
| ckpt_340k.pt | 340000 | 47.01 | 75.89 | 2.3798 | 75.49 |
|
| 10 |
+
| ckpt_360k.pt | 360000 | 46.30 | 76.24 | 2.3765 | 75.37 |
|
| 11 |
+
| ckpt_380k.pt | 380000 | 46.38 | 75.87 | 2.3784 | 75.27 |
|
| 12 |
+
| ckpt_400k.pt | 400000 | 46.80 | 75.57 | 2.3722 | 75.32 |
|
| 13 |
+
|
| 14 |
+
**Verdict:** seed 1338 minus seed 1337 on eff: +0.26 / β0.34 / β0.32 / β0.28 at 340k / 360k / 380k / 400k. The seed alone moves eff by about Β±0.3, ARC-Easy by 0.5-1.4 pp and BLiMP by 0.1-0.5. None of the data arms in rounds 3-4 differ from each other by more than this.
|
| 15 |
+
|
| 16 |
+
## Common setup
|
| 17 |
+
|
| 18 |
+
- **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs. **Exception for this arm:** the data is identical to `r3B_arcmix_qa2x/` and only the random seed differs (1338 vs 1337).
|
| 19 |
+
- **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for rounds 3-4) with the same harness; the difference between arms is attributed to the data.
|
| 20 |
+
- **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
|
| 21 |
+
- **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed (`r4B_arcmix_qa2x_seed1338/` vs `r3B_arcmix_qa2x/`) moved eff by +0.26 / β0.34 / β0.32 / β0.28 at 340k-400k, so run-to-run noise is about 0.3 eff (one seed pair: an order of magnitude, not a precise estimate).
|
| 22 |
+
- **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise.
|
| 23 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
|
| 24 |
+
- **Data:** ARC-MIX, see the root card of this repository.
|
r4K_anchor_arcmix_qa1/README.md
ADDED
|
@@ -0,0 +1,24 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# r4K_anchor_arcmix_qa1 β round 4: anchor arm (ARC-MIX, no upweighting)
|
| 2 |
+
|
| 3 |
+
**Why it exists:** the baseline for round 3. Same pool size and source file as `r3B_arcmix_qa2x/` (ARC-MIX sampled uniformly, 2.6B-token pool) but with Q&A documents taken once, at their natural share (3.7%). Comparing r3B against this arm isolates the effect of the 2Γ Q&A upweight; comparing `r3A_arcmix_edu_clean/` against it isolates the effect of the educational data.
|
| 4 |
+
|
| 5 |
+
## Results
|
| 6 |
+
|
| 7 |
+
| checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
|
| 8 |
+
|---|---|---|---|---|---|
|
| 9 |
+
| ckpt_340k.pt | 340000 | 47.10 | 76.10 | 2.3736 | 75.61 |
|
| 10 |
+
| ckpt_360k.pt | 360000 | 47.60 | 75.60 | 2.3736 | 75.61 |
|
| 11 |
+
| ckpt_380k.pt | 380000 | 46.89 | 76.08 | 2.3696 | 75.54 |
|
| 12 |
+
| ckpt_400k.pt | 400000 | 47.73 | 75.25 | 2.3732 | 75.53 |
|
| 13 |
+
|
| 14 |
+
**Verdict:** Q&A Γ2 minus anchor on eff: β0.38 / +0.10 / +0.05 / +0.07; educational data minus anchor: +0.08 / β0.03 / +0.03 / +0.18 (340k-400k). Both are within the seed noise of about 0.3 eff (`r4B_arcmix_qa2x_seed1338/`). The only repeatable effect is that the educational data makes WikiText-2 slightly worse (by 0.003-0.010 byte-ppl) at all four checkpoints. The anchor itself ends 0.28 eff below the flagship; no arm of rounds 3-4 beats it.
|
| 15 |
+
|
| 16 |
+
## Common setup
|
| 17 |
+
|
| 18 |
+
- **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs.
|
| 19 |
+
- **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for rounds 3-4) with the same harness; the difference between arms is attributed to the data.
|
| 20 |
+
- **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
|
| 21 |
+
- **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed (`r4B_arcmix_qa2x_seed1338/` vs `r3B_arcmix_qa2x/`) moved eff by +0.26 / β0.34 / β0.32 / β0.28 at 340k-400k, so run-to-run noise is about 0.3 eff (one seed pair: an order of magnitude, not a precise estimate).
|
| 22 |
+
- **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise.
|
| 23 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
|
| 24 |
+
- **Data:** ARC-MIX, see the root card of this repository.
|