Model card update (BLiMP recomputed after harness fix; data-attribution experiments) + READMEs for 4 experiment folders
Browse files- README.md +168 -154
- ctrl_arcmix_resume/README.md +23 -0
- forkB_arcmix_edu/README.md +26 -0
- r3A_arcmix_edu_clean/README.md +23 -0
- r3B_arcmix_qa2x/README.md +23 -0
README.md
CHANGED
|
@@ -1,154 +1,168 @@
|
|
| 1 |
-
---
|
| 2 |
-
license: apache-2.0
|
| 3 |
-
language:
|
| 4 |
-
- en
|
| 5 |
-
library_name: pytorch
|
| 6 |
-
pipeline_tag: text-generation
|
| 7 |
-
tags:
|
| 8 |
-
- tiny-lm
|
| 9 |
-
- gpt
|
| 10 |
-
- nanogpt
|
| 11 |
-
- glint-tiny-ml-leaderboard
|
| 12 |
-
- english
|
| 13 |
-
datasets:
|
| 14 |
-
- SlayerLab/minimal-en-corpus-5b
|
| 15 |
-
---
|
| 16 |
-
|
| 17 |
-
# GoLLeM-v5 β Tiny English Language Models (16M-64M)
|
| 18 |
-
|
| 19 |
-
Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
|
| 20 |
-
(nanoGPT lineage) trained for the
|
| 21 |
-
[Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
## Model details
|
| 26 |
-
|
| 27 |
-
- **Architecture:** 16M/32M = decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings. **64M flagship = Qwen3-style decoder** (RoPE ΞΈ=100k, SwiGLU, RMSNorm, QK-Norm, value residuals)
|
| 28 |
-
- **Sizes:** 16M = 6 layers / d_model 408 / 6 heads (17.4M); 32M = 6 layers / d_model 576 / 9 heads (31.
|
| 29 |
-
- **Context length:** 1024 tokens.
|
| 30 |
-
- **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
|
| 31 |
-
- **Training:** AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024
|
| 32 |
-
|
| 33 |
-
## Checkpoints
|
| 34 |
-
|
| 35 |
-
| checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
|
| 36 |
-
|---|---|---|---|---|---|---|
|
| 37 |
-
| `bpe16m_3.2B/ckpt.pt` | 17.4M | L6 d408 h6 | 3.2B | 67.40 | 38.22 | 1.2161 |
|
| 38 |
-
| `bpe16m_6B/ckpt.pt` | 17.4M | L6 d408 h6 | 6B | 68.92 | 39.10 | 1.1943 |
|
| 39 |
-
| `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 |
|
| 40 |
-
| `bpe32m_baseline/ckpt.pt` | 31.
|
| 41 |
-
| `run_16m_expanded/ckpt.pt` (
|
| 42 |
-
| `run_32m_16b/ckpt.pt` (
|
| 43 |
-
| `run_149m/ckpt.pt` (scaling ref) | 149M | β | 10BΒ§ | 76.99 | 49.66 | 1.2052 |
|
| 44 |
-
| `run_32m_18b/ckpt.pt` (
|
| 45 |
-
| `
|
| 46 |
-
| `v1_muon/ckpt_400k.pt` (**flagship**
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
``
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
|
| 99 |
-
```
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
#
|
| 115 |
-
|
| 116 |
-
|
| 117 |
-
|
| 118 |
-
-
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
- **
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
|
| 135 |
-
|
| 136 |
-
|
| 137 |
-
|
| 138 |
-
|
| 139 |
-
|
| 140 |
-
|
| 141 |
-
-
|
| 142 |
-
|
| 143 |
-
|
| 144 |
-
|
| 145 |
-
|
| 146 |
-
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
|
| 150 |
-
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
library_name: pytorch
|
| 6 |
+
pipeline_tag: text-generation
|
| 7 |
+
tags:
|
| 8 |
+
- tiny-lm
|
| 9 |
+
- gpt
|
| 10 |
+
- nanogpt
|
| 11 |
+
- glint-tiny-ml-leaderboard
|
| 12 |
+
- english
|
| 13 |
+
datasets:
|
| 14 |
+
- SlayerLab/minimal-en-corpus-5b
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# GoLLeM-v5 β Tiny English Language Models (16M-64M)
|
| 18 |
+
|
| 19 |
+
Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
|
| 20 |
+
(nanoGPT lineage) trained for the
|
| 21 |
+
[Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
|
| 22 |
+
The repository holds a controlled scaling study (tokens, width, optimizer) and a set of
|
| 23 |
+
**data-attribution experiments** on the 64M model (paired continued-training runs that differ only in data).
|
| 24 |
+
|
| 25 |
+
## Model details
|
| 26 |
+
|
| 27 |
+
- **Architecture:** 16M/32M = decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings. **64M flagship = Qwen3-style decoder** (RoPE ΞΈ=100k, SwiGLU, RMSNorm, QK-Norm, value residuals).
|
| 28 |
+
- **Sizes:** 16M = 6 layers / d_model 408 / 6 heads (17.4M); 32M = 6 layers / d_model 576 / 9 heads (31.6M); **64M flagship = 14 layers / d_model 576 / 9 heads (62.9M)**.
|
| 29 |
+
- **Context length:** 1024 tokens.
|
| 30 |
+
- **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
|
| 31 |
+
- **Training:** 16M/32M: AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024, seed 1337, bf16 (RTX 5090). 64M flagship: **Muon** (muon-lr 0.02) + AdamW for non-matrix parameters, same cosine schedule, 400k steps.
|
| 32 |
+
|
| 33 |
+
## Checkpoints
|
| 34 |
+
|
| 35 |
+
| checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
|
| 36 |
+
|---|---|---|---|---|---|---|
|
| 37 |
+
| `bpe16m_3.2B/ckpt.pt` | 17.4M | L6 d408 h6 | 3.2B | 67.40βΊ | 38.22 | 1.2161 |
|
| 38 |
+
| `bpe16m_6B/ckpt.pt` | 17.4M | L6 d408 h6 | 6B | 68.92βΊ | 39.10 | 1.1943 |
|
| 39 |
+
| `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36Β° | 39.52 | 1.1815 |
|
| 40 |
+
| `bpe32m_baseline/ckpt.pt` | 31.6M | L6 d576 h9 | 10B | 70.08Λ’ | 42.59 | 1.124 |
|
| 41 |
+
| `run_16m_expanded/ckpt.pt` (16M board entry) | 17.4M | L6 d408 h6 | 16Bβ | **70.08** | 40.91 | 1.4193 |
|
| 42 |
+
| `run_32m_16b/ckpt.pt` (32M board entry) | 31.6M | L6 d576 h9 | 16Bβ‘ | **73.48** | **44.44** | 1.3441 |
|
| 43 |
+
| `run_149m/ckpt.pt` (scaling ref) | 149M | β | 10BΒ§ | 76.99Β° | 49.66 | 1.2052 |
|
| 44 |
+
| `run_32m_18b/ckpt.pt` (slope-check) | 31.6M | L6 d576 h9 | 18BΒΆ | 72.38Β° | 44.70 | 1.3431 |
|
| 45 |
+
| Muon 32M (results only: `glint_32m_muon_results.json`, checkpoint not published) | 31.6M | L6 d576 h9 | 16Bβ | 72.29Β° | 42.89 | 1.3866 |
|
| 46 |
+
| `v1_muon/ckpt_400k.pt` (**64M flagship**) | 62.9M | L14 d576 h9 | 13.1Bβ
| **75.83** | **47.94** | 1.246 |
|
| 47 |
+
|
| 48 |
+
**BLiMP harness fix (2026-09-25).** Our evaluation script matched BLiMP configs by substring, so 6 phenomena were counted twice (73,000 pairs instead of 67,000). It is fixed in `glint_parity_eval.py` (exact config match + assertion of 67,000 pairs). The three board entries (bold BLiMP) were recomputed with the fixed script; for each we also reproduced the old 73,000-pair variant, which matches the previously published number to two decimals, so the difference comes from the loader alone (+0.16 to +0.46 pp before the fix). Other rows are **earlier results, not recomputed**:
|
| 49 |
+
- Β° computed with the pre-fix 73,000-pair loader (pair count recorded in the results file); on the three recomputed models the bias was +0.16 to +0.46 pp upward, expect similar here;
|
| 50 |
+
- Λ’ BLiMP on a 4,000-pair random sample, so the error is random (SE ~0.7 pp), not a systematic bias;
|
| 51 |
+
- βΊ the results file is not archived, so the pair count is unknown; treat as indicative only.
|
| 52 |
+
|
| 53 |
+
β 16M board entry = expanded 8.29B-token corpus (~1.9 epochs). eff **74.33** with the fixed BLiMP (70.08 / 40.91 / byte_ppl 2.6746).
|
| 54 |
+
|
| 55 |
+
β‘ 32M board entry = 32M at 16B tokens on the ARC-MIX corpus. eff **75.41** with the fixed BLiMP (73.48 / 44.44 / byte_ppl 2.5386). It broke the 16M BLiMP ceiling (~70) seen in the token scan: more capacity plus tokens moved both axes.
|
| 56 |
+
|
| 57 |
+
Β§ 149M = scaling reference only (under-trained at 67 tok/param). Highest raw scores in the older runs, but the size multiplier of the efficiency score falls with size, so it ranks lower on eff.
|
| 58 |
+
|
| 59 |
+
ΒΆ Slope-check = 32M at 18B tokens on the same ARC-MIX corpus as the 32M entry. BLiMP 72.38 (lower than at 16B) with byte_ppl flat: more epochs over a fixed corpus did not help. Diagnostic run.
|
| 60 |
+
|
| 61 |
+
β 32M with Muon on the expanded corpus. This run changed optimizer and corpus at once, so it cannot isolate Muon. The clean optimizer comparison was later run at 64M (see β
). Diagnostic run.
|
| 62 |
+
|
| 63 |
+
β
64M flagship (v1 Muon) = Qwen3-style decoder + value residuals, **Muon** optimizer, ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). Recomputed with the fixed harness: **ARC-Easy 47.94 / BLiMP 75.83 / WikiText-2 byte_ppl 2.372 (BPB 1.246) β eff 75.81.** The maintainer's independent re-benchmark (PR #78, sha256 `59f982c1β¦` matched) reproduced ARC-Easy 47.94 exactly; its BLiMP 75.99 matches our 73,000-pair variant to two decimals, so it was most likely computed with the same (pre-fix) script. Earlier versions of this card cited eff 77.51 (byte_ppl 2.016, from a wrong bytes-per-token factor 4.755 instead of 3.8605) and later eff ~75.9 (BLiMP 75.99); both are superseded. Muon was at least as good as AdamW in a clean 64M A/B (identical data, architecture and seed; only the optimizer differs); those numbers pre-date the BLiMP fix.
|
| 64 |
+
|
| 65 |
+
## Data-attribution experiments (64M, 2026-09-25)
|
| 66 |
+
|
| 67 |
+
Question: **does changing the training data in the late phase of training move the efficiency score?** Method: two identical 64M models continue from the same flagship checkpoint (step 320k) with the same recipe, seed and number of steps (80k, to step 400k); **they differ only in the data**. Both are evaluated at 340k / 360k / 380k / 400k with the board protocol and a paired bootstrap over items. Each checkpoint folder has its own README with the full data specification.
|
| 68 |
+
|
| 69 |
+
| run (folder) | data | ARC-E | BLiMP | wiki byte_ppl | eff @400k | note |
|
| 70 |
+
|---|---|---|---|---|---|---|
|
| 71 |
+
| flagship (`v1_muon/`) | ARC-MIX, full run 0β400k | 47.94 | 75.83 | 2.3718 | 75.81 | reference |
|
| 72 |
+
| control (`ctrl_arcmix_resume/`) | ARC-MIX | 47.69 | 76.29 | 2.3717 | 75.88 | re-draws the flagship's first 80k steps of data (the data sampler restarted from its seed on resume) |
|
| 73 |
+
| fork-B (`forkB_arcmix_edu/`) | ARC-MIX + educational | 46.34 | 77.20 | 2.3982 | 75.66 | **flawed build**: took only the head of the ARC-MIX file and a different document separator. Not a test of educational data |
|
| 74 |
+
| A (`r3A_arcmix_edu_clean/`) | 45% ARC-MIX + 55% educational (clean build), 2.55B-token pool | 47.31 | 76.22 | 2.3762 | 75.71 | clean pair with B |
|
| 75 |
+
| B (`r3B_arcmix_qa2x/`) | ARC-MIX with Q&A documents Γ2, 2.62B-token pool | 47.26 | 75.90 | 2.3712 | 75.60 | clean pair with A |
|
| 76 |
+
|
| 77 |
+
- **A β B: Ξeff +0.11 (95% CI β0.30 to +0.51)**; at the four checkpoints +0.46 / β0.13 / β0.02 / +0.11. No difference on eff. Educational data consistently made WikiText-2 slightly worse (by more than the run-to-run spread); BLiMP was slightly higher but within noise.
|
| 78 |
+
- **Conclusion:** at these data doses and 80k steps of the cosine tail (learning rate 19% β 10% of peak), the effect of the data on eff is **below ~0.5**. None of the runs beats the flagship beyond noise, so none is submitted to the board. This is an upper bound on the effect at this scale of intervention, not evidence that data does not matter.
|
| 79 |
+
- **In progress:** an anchor run (ARC-MIX sample of the same pool size, Q&A Γ1) and B with a second seed, to measure the Q&A effect cleanly and the run-to-run noise directly.
|
| 80 |
+
|
| 81 |
+
## Usage
|
| 82 |
+
|
| 83 |
+
These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
|
| 84 |
+
weights. The model class and a board-scoring harness are included in this repo:
|
| 85 |
+
|
| 86 |
+
- `train_gpt_ref.py` β GPT definition (rebuild the GPT of the tabled shape, `load_state_dict`, trim logits to vocab 12288).
|
| 87 |
+
- `glint_parity_eval.py` β Glint board-scoring forward (256-token clip, raw log-prob) for BLiMP / ARC-Easy / WikiText-2 (BLiMP loader fixed 2026-09-25, see above).
|
| 88 |
+
|
| 89 |
+
```python
|
| 90 |
+
import torch
|
| 91 |
+
from tokenizers import Tokenizer
|
| 92 |
+
tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
|
| 93 |
+
ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
|
| 94 |
+
state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
|
| 95 |
+
```
|
| 96 |
+
|
| 97 |
+
**64M checkpoints (`v1_muon/`, and the data-attribution folders) β Qwen3-style arch, config embedded in the checkpoint.** Load with the cfg from `ckpt["config"]`, **not** the legacy 5-arg nanoGPT path. `glint_parity_eval.py`'s loader is cfg-aware for both lineages.
|
| 98 |
+
|
| 99 |
+
```python
|
| 100 |
+
import torch
|
| 101 |
+
from types import SimpleNamespace
|
| 102 |
+
from train_gpt_ref import GPT
|
| 103 |
+
ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
|
| 104 |
+
c = ck["config"] # self-describing: vocab 12288, n_layer 14, n_embd 576, n_head 9, block 1024
|
| 105 |
+
cfg = SimpleNamespace(**c) # Qwen3 flags: rope ΞΈ100k / swiglu / qk_norm / value_residual
|
| 106 |
+
m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
|
| 107 |
+
m.load_state_dict(ck["model"], strict=False) # tied head.weight
|
| 108 |
+
m.eval()
|
| 109 |
+
```
|
| 110 |
+
|
| 111 |
+
**Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults:
|
| 112 |
+
|
| 113 |
+
```bash
|
| 114 |
+
# 16M @ expanded corpus
|
| 115 |
+
python train_gpt_ref.py --data-dir <corpus> \
|
| 116 |
+
--n-layer 6 --n-embd 408 --n-head 6 --block 1024 \
|
| 117 |
+
--batch 64 --steps 244141 --lr 6e-4 --min-lr 6e-5 \
|
| 118 |
+
--vocab 12288 --dtype uint16 --seed 1337
|
| 119 |
+
# 32M: --n-embd 576 --n-head 9 (same vocab 12288 / uint16 / BPE tokenizer)
|
| 120 |
+
```
|
| 121 |
+
|
| 122 |
+
The script's byte-level defaults (`--vocab 256 --dtype uint8`) and its header comment reflect its origin as a **standard-GPT control** compared against an experimental **BDH (fast-weights)** architecture; the leaderboard models here are the standard causal transformer in BPE mode and do **not** use BDH.
|
| 123 |
+
|
| 124 |
+
## Training data
|
| 125 |
+
|
| 126 |
+
[`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
|
| 127 |
+
β ~5.40B BPE-12k tokens, English, decontaminated against the benchmark test sets. A broad high-quality mix:
|
| 128 |
+
FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
|
| 129 |
+
A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) feeds later runs.
|
| 130 |
+
|
| 131 |
+
**ARC-MIX (9.42B).** The 32M board entry and the 64M runs use **ARC-MIX**: a reasoning/knowledge-enriched expansion of the base mixture (9,417,035,832 BPE-12288 tokens) with ARC-relevant science/reasoning/QA web content upweighted (gold ~3Γ, related ~2Γ) over the decontaminated base. Same tokenizer, same 13-gram benchmark decontamination (WikiText-2 / BLiMP / ARC test sets).
|
| 132 |
+
|
| 133 |
+
**Earlier FineWeb-Edu-dominant corpus (v2).** A from-scratch 64M run on a FineWeb-Edu-dominant blend scored lower on ARC than ARC-MIX. That build used a different document separator than the rest of our data and a changed recipe (WSD schedule, z-loss, logit cap), so the comparison is confounded and is **not** evidence against educational data trained from scratch.
|
| 134 |
+
|
| 135 |
+
## Evaluation
|
| 136 |
+
|
| 137 |
+
All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e. the board-comparable definitions:
|
| 138 |
+
|
| 139 |
+
- **BLiMP** β 67 configs (train split, 67,000 pairs), each sentence clipped to the first 256 tokens, raw sentence log-prob preference (`good > bad`), no length normalization.
|
| 140 |
+
- **ARC-Easy** β test split, zero-shot, raw accuracy over `LL(question + choice) - LL(question)`.
|
| 141 |
+
- **WikiText-2** β byte-normalized (byte_ppl / BPB; bytes-per-token 3.8605 for this tokenizer on WikiText-2-raw-v1 test).
|
| 142 |
+
|
| 143 |
+
A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
|
| 144 |
+
|
| 145 |
+
**Board status.** The Glint board still lists the earlier values for our entries (64M: BLiMP 77.84 / byte_ppl 2.016; 32M: BLiMP 73.77; 16M: BLiMP 70.53). The correct values are the ones in this card. With the board's own efficiency formula (`eff = mean(BLiMP, ARC-Easy, normalized-WikiText-2) Γ size multiplier`, checked line-for-line against the Space source) they give eff **75.81 (64M), 75.41 (32M), 74.33 (16M)**.
|
| 146 |
+
|
| 147 |
+
## Key findings
|
| 148 |
+
|
| 149 |
+
- **Tokens drive BLiMP at small size, up to a size-specific ceiling.** At 16M, BLiMP rose with tokens (3.2B β 10B) and then flattened near ~70 on the expanded corpus; 32M at 16B tokens moved past that ceiling.
|
| 150 |
+
- **Capacity and knowledge drive ARC.** 16M β 32M at matched tokens lifted ARC-Easy by ~3pp; 64M lifted it further (47.94).
|
| 151 |
+
- **Bigger is not automatically better on eff.** Raw scores keep rising with size, but the efficiency score multiplies by a size bonus that shrinks with parameter count. In our runs 64M is the best eff entry (75.81 vs 75.41 at 32M); a 128M run improves raw quality but the smaller multiplier cancels most of it.
|
| 152 |
+
- **Muon was at least as good as AdamW at 64M** in a clean A/B (same data, architecture, seed; numbers pre-date the BLiMP fix).
|
| 153 |
+
- **Late-phase data changes move eff by less than ~0.5** (see Data-attribution experiments).
|
| 154 |
+
- **Value residuals** are part of the 64M stack. The board's #1 model attributes ~+6 ARC-Easy to them in its own card; we have **not** isolated this effect ourselves.
|
| 155 |
+
|
| 156 |
+
## Roadmap
|
| 157 |
+
|
| 158 |
+
- Continue training the 64M flagship in segments with a constant-learning-rate phase and a final decay, comparing two data mixes per segment on **validation** sets (never the board test sets) and keeping the winner only when it wins beyond measured run-to-run noise.
|
| 159 |
+
- A clean from-scratch data comparison (same recipe and separator) and synthetic science data for ARC are candidates; both are pending.
|
| 160 |
+
|
| 161 |
+
## Limitations
|
| 162 |
+
|
| 163 |
+
Base (not instruction-tuned) research models at 16-64M parameters, English-only. Expect limited factual knowledge and
|
| 164 |
+
coherence; not intended for production use.
|
| 165 |
+
|
| 166 |
+
## Provenance
|
| 167 |
+
|
| 168 |
+
Trained on RunPod RTX 5090. Full evaluation artifacts and protocol details are available on request.
|
ctrl_arcmix_resume/README.md
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# ctrl_arcmix_resume β control arm (ARC-MIX, unchanged data)
|
| 2 |
+
|
| 3 |
+
**Why it exists:** the reference arm of the first data A/B round. It continues the flagship on the same ARC-MIX corpus the flagship was trained on, so that the other arm (`forkB_arcmix_edu/`) can be compared against "more of the same data".
|
| 4 |
+
|
| 5 |
+
**Data:** ARC-MIX 9.42B-token corpus (the flagship's training data), unchanged.
|
| 6 |
+
|
| 7 |
+
**Important caveat:** when this run was resumed from step 320,000, the trainer restarted its batch sampler from the seed. Because the corpus file is identical to the flagship's, this arm re-drew exactly the same training windows the flagship saw in its first 80k steps. It is therefore a *replay* of early data, not fresh data, and comparisons against it carry that confound. A trainer fix (fast-forwarding the sampler on resume) is prepared for future runs.
|
| 8 |
+
|
| 9 |
+
## Results
|
| 10 |
+
|
| 11 |
+
| checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
|
| 12 |
+
|---|---|---|---|---|---|
|
| 13 |
+
| ckpt_360k.pt | 360000 | 46.63 | 75.76 | 2.3727 | 75.33 |
|
| 14 |
+
| ckpt_400k.pt | 400000 | 47.69 | 76.29 | 2.3717 | 75.88 |
|
| 15 |
+
|
| 16 |
+
## Common setup
|
| 17 |
+
|
| 18 |
+
- **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs.
|
| 19 |
+
- **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data.
|
| 20 |
+
- **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
|
| 21 |
+
- **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate).
|
| 22 |
+
- **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise.
|
| 23 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
|
forkB_arcmix_edu/README.md
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# forkB_arcmix_edu β first data arm (ARC-MIX + educational web text), confounded build
|
| 2 |
+
|
| 3 |
+
**Why it exists:** the treatment arm of the first data A/B round: does adding educational web text (FineWeb-Edu, decontaminated against the evaluation test sets) to ARC-MIX in the last 80k steps improve the model? Mix: about 45% ARC-MIX, 55% educational text, 2.55B-token pool.
|
| 4 |
+
|
| 5 |
+
**Important caveat β this is not a clean test of educational data.** The data build had two defects found during training, before evaluation:
|
| 6 |
+
1. the ARC-MIX half was taken from the beginning of an unshuffled, source-ordered file (about 1.15B tokens) instead of being sampled across the whole corpus, so it contains almost none of the chat/Q&A-format documents present in ARC-MIX;
|
| 7 |
+
2. educational documents were separated by `<|im_end|>` instead of the corpus end-of-text token.
|
| 8 |
+
The result below is therefore reported only as "this blend vs control". The clean rebuild is `r3A_arcmix_edu_clean/`.
|
| 9 |
+
|
| 10 |
+
## Results
|
| 11 |
+
|
| 12 |
+
| checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
|
| 13 |
+
|---|---|---|---|---|---|
|
| 14 |
+
| ckpt_360k.pt | 360000 | 45.83 | 76.81 | 2.3996 | 75.35 |
|
| 15 |
+
| ckpt_400k.pt | 400000 | 46.34 | 77.20 | 2.3982 | 75.66 |
|
| 16 |
+
|
| 17 |
+
Versus the control arm (`ctrl_arcmix_resume/`), paired bootstrap on eff: +0.02 [β0.41, +0.45] at 360k, β0.22 [β0.64, +0.21] at 400k β no difference. Consistent pattern: BLiMP higher, WikiText-2 byte-perplexity worse. The clean rebuild shows this pattern came from the build defects, not from the educational data.
|
| 18 |
+
|
| 19 |
+
## Common setup
|
| 20 |
+
|
| 21 |
+
- **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs.
|
| 22 |
+
- **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data.
|
| 23 |
+
- **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
|
| 24 |
+
- **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate).
|
| 25 |
+
- **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise.
|
| 26 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
|
r3A_arcmix_edu_clean/README.md
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# r3A_arcmix_edu_clean β arm A of round 3: ARC-MIX + educational text, clean build
|
| 2 |
+
|
| 3 |
+
**Why it exists:** a clean repeat of the first educational-data arm. Same proportion as `forkB_arcmix_edu/` (45.1% ARC-MIX, 54.9% educational text, 2.55B-token pool), with the build fixed: ARC-MIX sampled uniformly across the whole corpus at document boundaries, and every document separated by the corpus end-of-text token. It was trained in parallel with arm B (`r3B_arcmix_qa2x/`) from the same checkpoint.
|
| 4 |
+
|
| 5 |
+
## Results
|
| 6 |
+
|
| 7 |
+
| checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
|
| 8 |
+
|---|---|---|---|---|---|
|
| 9 |
+
| ckpt_340k.pt | 340000 | 47.35 | 76.13 | 2.3801 | 75.69 |
|
| 10 |
+
| ckpt_360k.pt | 360000 | 47.18 | 76.01 | 2.3839 | 75.58 |
|
| 11 |
+
| ckpt_380k.pt | 380000 | 47.10 | 76.04 | 2.3787 | 75.57 |
|
| 12 |
+
| ckpt_400k.pt | 400000 | 47.31 | 76.22 | 2.3762 | 75.71 |
|
| 13 |
+
|
| 14 |
+
**Verdict (A β B, paired bootstrap on eff):** +0.46 / β0.13 / β0.02 / +0.11 at 340k / 360k / 380k / 400k; at 400k +0.11 [β0.30, +0.51]. **No difference on eff.** Across all four checkpoints, arm A has slightly worse WikiText-2 byte-perplexity (a consistent effect of the educational data at this dose) and slightly higher BLiMP (within noise). Conclusion recorded for the study: at this data dose, changing data only in the last 80k steps moves eff by less than about 0.5.
|
| 15 |
+
|
| 16 |
+
## Common setup
|
| 17 |
+
|
| 18 |
+
- **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs.
|
| 19 |
+
- **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data.
|
| 20 |
+
- **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
|
| 21 |
+
- **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate).
|
| 22 |
+
- **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise.
|
| 23 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
|
r3B_arcmix_qa2x/README.md
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# r3B_arcmix_qa2x β arm B of round 3: ARC-MIX with Q&A documents upweighted
|
| 2 |
+
|
| 3 |
+
**Why it exists:** tests whether upweighting question-answering data helps. ARC-MIX sampled uniformly across the whole corpus (2.62B-token pool), with documents in chat/Q&A format (those containing `<|im_start|>`) taken twice, raising their share from 3.7% to 7.1%. No educational data. Trained in parallel with arm A (`r3A_arcmix_edu_clean/`) from the same checkpoint.
|
| 4 |
+
|
| 5 |
+
## Results
|
| 6 |
+
|
| 7 |
+
| checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
|
| 8 |
+
|---|---|---|---|---|---|
|
| 9 |
+
| ckpt_340k.pt | 340000 | 46.42 | 75.71 | 2.3781 | 75.23 |
|
| 10 |
+
| ckpt_360k.pt | 360000 | 47.73 | 75.77 | 2.3744 | 75.71 |
|
| 11 |
+
| ckpt_380k.pt | 380000 | 47.39 | 75.77 | 2.3756 | 75.59 |
|
| 12 |
+
| ckpt_400k.pt | 400000 | 47.26 | 75.90 | 2.3712 | 75.60 |
|
| 13 |
+
|
| 14 |
+
**Verdict:** see `r3A_arcmix_edu_clean/` β A β B shows no difference on eff (400k: +0.11 [β0.30, +0.51]). A follow-up round compares this arm against an anchor with the same pool size and no upweighting (`r4K_anchor_arcmix_qa1/`) and against the same data with a different seed (`r4B_arcmix_qa2x_seed1338/`); first readings suggest the 2Γ Q&A upweight slightly hurts rather than helps, not yet beyond noise.
|
| 15 |
+
|
| 16 |
+
## Common setup
|
| 17 |
+
|
| 18 |
+
- **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs.
|
| 19 |
+
- **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data.
|
| 20 |
+
- **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
|
| 21 |
+
- **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate).
|
| 22 |
+
- **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise.
|
| 23 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
|