|
Download forkB_arcmix_edu/README.md from SlayerLab/gollem-v5-ckpts: direct link, hf CLI and curl.
- Browser
- Download file 3.11 kB
-
https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/0da09be998d4e11cb747939d94b2ec6a62f5a541/forkB_arcmix_edu/README.md
- Command line
-
hf download hf://SlayerLab/gollem-v5-ckpts@0da09be998d4e11cb747939d94b2ec6a62f5a541/forkB_arcmix_edu/README.md
-
curl -L -o README.md https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/0da09be998d4e11cb747939d94b2ec6a62f5a541/forkB_arcmix_edu/README.md
3.11 kB
forkB_arcmix_edu — first data arm (ARC-MIX + educational web text), confounded build
Why it exists: the treatment arm of the first data A/B round: does adding educational web text (FineWeb-Edu, decontaminated against the evaluation test sets) to ARC-MIX in the last 80k steps improve the model? Mix: about 45% ARC-MIX, 55% educational text, 2.55B-token pool.
Important caveat — this is not a clean test of educational data. The data build had two defects found during training, before evaluation:
- the ARC-MIX half was taken from the beginning of an unshuffled, source-ordered file (about 1.15B tokens) instead of being sampled across the whole corpus, so it contains almost none of the chat/Q&A-format documents present in ARC-MIX;
- educational documents were separated by
<|im_end|>instead of the corpus end-of-text token. The result below is therefore reported only as "this blend vs control". The clean rebuild isr3A_arcmix_edu_clean/.
Results
| checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
|---|---|---|---|---|---|
| ckpt_360k.pt | 360000 | 45.83 | 76.81 | 2.3996 | 75.35 |
| ckpt_400k.pt | 400000 | 46.34 | 77.20 | 2.3982 | 75.66 |
Versus the control arm (ctrl_arcmix_resume/), paired bootstrap on eff: +0.02 [−0.41, +0.45] at 360k, −0.22 [−0.64, +0.21] at 400k — no difference. Consistent pattern: BLiMP higher, WikiText-2 byte-perplexity worse. The clean rebuild shows this pattern came from the build defects, not from the educational data.
Common setup
- Base model: GoLLeM-v5 64M flagship (
v1_muon/, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs. - Method: two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data.
- Evaluation:
glint_parity_eval.pyin the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name. - Reference: flagship
v1_muon/ckpt_400k.ptscores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate). - Status: research checkpoint, not a leaderboard submission. No arm of this study beats the flagship beyond run-to-run noise.
- Format: PyTorch checkpoint dict with
model,opt,step,config;train_gpt_ref.pyin the repository root rebuilds the model fromconfig.