Maggio33 commited on
Commit
1cd1e4d
Β·
verified Β·
1 Parent(s): 3b05e58

Model card update (BLiMP recomputed after harness fix; data-attribution experiments) + READMEs for 4 experiment folders

Browse files
README.md CHANGED
@@ -1,154 +1,168 @@
1
- ---
2
- license: apache-2.0
3
- language:
4
- - en
5
- library_name: pytorch
6
- pipeline_tag: text-generation
7
- tags:
8
- - tiny-lm
9
- - gpt
10
- - nanogpt
11
- - glint-tiny-ml-leaderboard
12
- - english
13
- datasets:
14
- - SlayerLab/minimal-en-corpus-5b
15
- ---
16
-
17
- # GoLLeM-v5 β€” Tiny English Language Models (16M-64M)
18
-
19
- Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
20
- (nanoGPT lineage) trained for the
21
- [Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
22
- This repository is a controlled single-factor scaling study: identical architecture/hyper-parameters/seed,
23
- varying only tokens and model width.
24
-
25
- ## Model details
26
-
27
- - **Architecture:** 16M/32M = decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings. **64M flagship = Qwen3-style decoder** (RoPE ΞΈ=100k, SwiGLU, RMSNorm, QK-Norm, value residuals), trained with the **Muon** optimizer.
28
- - **Sizes:** 16M = 6 layers / d_model 408 / 6 heads (17.4M); 32M = 6 layers / d_model 576 / 9 heads (31.4M); **64M flagship = 14 layers / d_model 576 / 9 heads (62.9M)**.
29
- - **Context length:** 1024 tokens.
30
- - **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
31
- - **Training:** AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024 (65,536 tok/step), seed 1337, bf16 (RTX 5090).
32
-
33
- ## Checkpoints
34
-
35
- | checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
36
- |---|---|---|---|---|---|---|
37
- | `bpe16m_3.2B/ckpt.pt` | 17.4M | L6 d408 h6 | 3.2B | 67.40 | 38.22 | 1.2161 |
38
- | `bpe16m_6B/ckpt.pt` | 17.4M | L6 d408 h6 | 6B | 68.92 | 39.10 | 1.1943 |
39
- | `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 |
40
- | `bpe32m_baseline/ckpt.pt` | 31.4M | L6 d576 h9 | 10B | 70.08 | 42.59 | 1.124 |
41
- | `run_16m_expanded/ckpt.pt` (crown) | 17.4M | L6 d408 h6 | 16B† | 70.53 | 40.91 | 1.4193 |
42
- | `run_32m_16b/ckpt.pt` (Path-B v1) | 31.4M | L6 d576 h9 | 16B‑ | **73.77** | **44.44** | 1.3441 |
43
- | `run_149m/ckpt.pt` (scaling ref) | 149M | β€” | 10BΒ§ | 76.99 | 49.66 | 1.2052 |
44
- | `run_32m_18b/ckpt.pt` (v1b slope-check) | 31.6M | L6 d576 h9 | 18BΒΆ | 72.38 | 44.70 | 1.3431 |
45
- | `run_32m_muon/ckpt.pt` (Muon optimizer) | 31.6M | L6 d576 h9 | 16Bβ€– | 72.29 | 42.89 | 1.3866 |
46
- | `v1_muon/ckpt_400k.pt` (**flagship** β€” Qwen3+Muon+VR) | 62.9M | L14 d576 h9 | 13.1Bβ˜… | 75.99 | **47.94** | 1.246 |
47
-
48
- † crown = expanded 8.29B-token corpus (~1.9 epochs). This is the published **16M board entry: confirmed #21** (eff 74.49 β€” 70.53 / 40.91 / byte_ppl 2.6746). Earlier recon estimated #20 (an optimistic #14 used a wrong wiki estimate before the exact byte_ppl correction); the official merge landed #21.
49
-
50
- ‑ Path-B v1 = 32M at 16B tokens (ARC-MIX corpus): **breaks the 16M BLiMP ceiling** (70.5 -> 73.77) and lifts ARC to 44.44. This is the published **32M board entry: confirmed #16** (eff 75.51); recon estimated #15, official merge landed #16. The 70.5 cap was 16M-specific, not absolute; more capacity + tokens moves both axes.
51
-
52
- Β§ 149M = scaling reference only (heavily under-trained at 67 tok/param). Highest raw scores (BLiMP 76.99 / ARC 49.66) but board-recon **#20/74** β€” the efficiency size-bonus caps at ~32M, so bigger models score higher raw but rank lower on eff. The eff sweet-spot is ~32M; the lever toward the top is raw-score at 32M (architecture / optimizer / tokens), not more size.
53
-
54
- ΒΆ v1b slope-check = 32M at 18B tokens on the **same arcmix corpus** as Path-B v1. BLiMP 72.38 (βˆ’1.39 vs v1@16B) with byte_ppl flat (2.537 vs 2.539) β€” **the arcmix corpus is saturated at ~16B**: more epochs (~1.9) over-cycle and mildly hurt BLiMP. This pre-registered slope-check refutes "train longer" on a fixed corpus; the next gain needs **unique** data (broad web), not re-cycled epochs. Diagnostic run (ckpt on request).
55
-
56
- β€– Muon optimizer = 32M at 16B tokens, Muon optimizer (muon-lr 0.02), on the expanded 8.29B corpus. BLiMP 72.29 / ARC 42.89 / byte_ppl 2.615. **Optimizer verdict = inconclusive (confounded design):** this run also changed corpus (expanded vs Path-B's arcmix), so the βˆ’1.48 BLiMP mixes optimizer *and* data and cannot isolate Muon. A clean Muon single-factor is deferred to the 64M A/B (Muon vs AdamW, corpus held fixed). Diagnostic run.
57
-
58
- β˜… 64M flagship (v1 Muon) = 62.9M, **Qwen3-style decoder** (RoPE ΞΈ=100k + SwiGLU + RMSNorm + QK-Norm + value residuals), **Muon** optimizer (muon-lr 0.02, cosine), ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). **Maintainer re-benchmark (PR #78, RTX 5090, strict load, sha256 `59f982c1…` matched): ARC-Easy 47.94 β€” an exact match to our number, the checkpoint-load + log-likelihood-scoring control β€” BLiMP 75.99, WikiText-2 byte_ppl 2.372 / BPB 1.246 β†’ eff ~75.9.** An earlier version of this card cited eff 77.51 / byte_ppl 2.016 / BPB 1.012; that byte_ppl used a **wrong bytes-per-token conversion** (4.755 vs the canonical raw-text-bytes/tokens = 1,292,008 / 334,674 = 3.86 on WikiText-2-raw-v1 test), now corrected. Board-protocol (Glint bare-prompt ARC β€” `LL(choice|q)` argmax, full BLiMP-67k, wiki byte-normalized, no-BOS). A clean single-factor scale-up of the 32M Path-B recipe (same BPE-12288 tokenizer / arcmix data lineage; only size + the Qwen3+Muon+VR arch differ). Also settles the clean **Muon verdict**: at 64M with corpus held fixed (Muon vs AdamW A/B), Muon wins on byte_ppl + BLiMP.
59
-
60
- ## Usage
61
-
62
- These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
63
- weights. The model class and a ready board-scoring harness are included in this repo:
64
-
65
- - `train_gpt_ref.py` β€” GPT definition (rebuild the GPT of the tabled shape, `load_state_dict`, trim logits to vocab 12288).
66
- - `glint_parity_eval.py` β€” the exact Glint board-scoring forward (256-token clip, raw log-prob) for BLiMP / ARC-Easy / WikiText-2.
67
-
68
- ```python
69
- import torch
70
- from tokenizers import Tokenizer
71
- tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
72
- ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
73
- state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
74
- ```
75
-
76
- **64M flagship (`v1_muon/ckpt_400k.pt`) β€” Qwen3-style arch, config embedded in the checkpoint.** Load with the cfg from `ckpt["config"]` (RoPE / SwiGLU / QK-Norm / value-residual, `n_head=9` / L14 / d576 / block 1024 are self-describing), **not** the legacy 5-arg nanoGPT path. `glint_parity_eval.py`'s loader is cfg-aware for both lineages.
77
-
78
- ```python
79
- import torch
80
- from types import SimpleNamespace
81
- from train_gpt_ref import GPT
82
- ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
83
- c = ck["config"] # self-describing: vocab 12288, n_layer 14, n_embd 576, n_head 9, block 1024
84
- cfg = SimpleNamespace(**c) # Qwen3 flags: rope ΞΈ100k / swiglu / qk_norm / value_residual
85
- m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
86
- m.load_state_dict(ck["model"], strict=False) # tied head.weight
87
- m.eval()
88
- ```
89
-
90
- **Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above):
91
-
92
- ```bash
93
- # crown 16M @ expanded corpus
94
- python train_gpt_ref.py --data-dir <corpus> \
95
- --n-layer 6 --n-embd 408 --n-head 6 --block 1024 \
96
- --batch 64 --steps 244141 --lr 6e-4 --min-lr 6e-5 \
97
- --vocab 12288 --dtype uint16 --seed 1337
98
- # Path-B 32M: --n-embd 576 --n-head 9 (same vocab 12288 / uint16 / BPE tokenizer)
99
- ```
100
-
101
- The `--vocab 12288 --dtype uint16` flags select BPE-12k over uint16 token bins. The script's byte-level defaults (`--vocab 256 --dtype uint8`) and its header comment reflect its origin as a **standard-GPT control** compared against an experimental **BDH (fast-weights)** architecture β€” the leaderboard models here are the standard causal transformer in BPE mode and do **not** use BDH.
102
-
103
- ## Training data
104
-
105
- [`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
106
- β€” ~5.40B BPE-12k tokens, English, decontaminated against the benchmark test sets. A broad high-quality mix:
107
- FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
108
- A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) feeds later runs.
109
-
110
- **ARC-MIX (9.42B).** The 32M Path-B and 64M A/B runs use **ARC-MIX** β€” a reasoning/knowledge-enriched expansion of the base mixture (9,417,035,832 BPE-12288 tokens): ARC-relevant science/reasoning/QA web content upweighted (gold ~3Γ—, related ~2Γ—) over the decontaminated base, to push the ARC-Easy axis (the binding efficiency constraint at this scale; capacity-gated per finding W11). Same BPE-12288 tokenizer, same 13-gram benchmark decontamination (WikiText-2 / BLiMP / ARC). The v2 #1-shot corpus moves to a FineWeb-Edu-dominant blend (β‰₯60% FineWeb-Edu + DCLM-baseline + FineMath-4plus, ~20B unique), per the 8M data-screen (FineWeb-Edu won BLiMP) and top-3 competitor recipes.
111
-
112
- **Training budget and epochs.** 16M trained on 16B tokens *seen* is intentional, not a chart error. The leaderboard scores efficiency = quality at a fixed tiny size, so you over-train to squeeze max quality from frozen capacity. Chinchilla-optimal (about 20x params = 0.32B for 16M) minimizes compute-optimal loss, which is NOT the leaderboard objective; top models train many tokens-per-param too. 16B seen over 8.29B unique corpus = about 1.9 epochs (each token seen ~1.9x, under the 2x repeat-degradation limit; val-loss healthy, zero memorization). Note: the crown 16B point changed BOTH tokens and corpus (5.4B to 8.29B expanded), so on the token-scan chart it is marked separately (star + dashed) and is not a pure token step. BLiMP saturation holds regardless: crown BLiMP 70.53 is below even the token-only projection 71.6.
113
-
114
- ## Evaluation
115
-
116
- All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e. the board-comparable definitions:
117
-
118
- - **BLiMP** β€” 67 configs (train split), each sentence clipped to the first 256 tokens, raw sentence log-prob preference (`good > bad`), no length normalization.
119
- - **ARC-Easy** β€” test split, zero-shot, raw accuracy over `LL(question + choice) - LL(question)`.
120
- - **WikiText-2** β€” byte-normalized bits-per-byte (the board's `wiki` field is byte-scale; token-perplexity is tokenizer-dependent and not directly comparable across models).
121
-
122
- A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
123
-
124
- **GoLLeM-v5 64M flagship (eff ~75.9), maintainer re-benchmarked (PR #78, 2026-09-24).** The 64M Muon model (Qwen3 arch + value residuals, ARC-MIX 9.42B) posts **ARC-Easy 47.94 β€” an exact match to our own number in the maintainer's independent RTX-5090 re-benchmark** (the checkpoint-load + log-likelihood-scoring control), with BLiMP 75.99 / WikiText-2 byte_ppl 2.372 β†’ eff ~75.9. It is a strong **dense 64M entry**, a **#16 β†’ 64M scale-up** from the 32M on the same data/tokenizer lineage plus the Qwen3+Muon+value-residual stack. An earlier version of this card cited eff 77.51 (byte_ppl 2.016) from a wrong bytes-per-token conversion β€” corrected here; better an honest ~75.9 than an inflated 77.5. Final board rank pending the maintainer's placement.
125
-
126
- **Positioning β€” CONFIRMED, on the board.** PR #76 was merged into the Glint Tiny-ML Leaderboard (2026-09-23), maintainer-verified (checkpoints loaded directly; params confirmed: 32M = 31,601,664, 16M = 17,449,344 deduped tied-embeddings; architecture matches `train_gpt_ref.py`, standard nanoGPT BPE-12288). Official standings: **#16 GoLLeM-v5 32M** (eff 75.51 β€” BLiMP 73.77 / ARC-Easy 44.44 / WikiText-2 byte_ppl 2.5386, 16B tok) and **#21 GoLLeM-v5 16M** (eff 74.49 β€” BLiMP 70.53 / ARC 40.91 / byte_ppl 2.6746). The board efficiency formula was reverse-engineered and then confirmed **line-for-line against the Space source** (reproduces the displayed eff exactly, 3/3 checked models to 2 decimals): `eff = mean(BLiMP, ARC-Easy, normalized-WikiText-2) Γ— size-bonus`, where the size-bonus runs 1.0Γ— (largest on board) to 1.5Γ— (smallest) on a log-parameter scale. Our 32M carries a 1.065Γ— size-bonus vs the #1's 1.013Γ— β€” a ~5% efficiency edge at equal raw metrics.
127
-
128
- ## Key findings (single-factor study)
129
-
130
- - **Tokens drive BLiMP, not size β€” up to a ceiling.** BLiMP climbs with tokens (~+1.8pp per doubling, 3.2B->10B) then **saturates at 16M's ~70.5 ceiling** (crown 16B: 70.53, +0.17 over 10B β€” flat, below the token-only projection 71.6); 16M->32M at matched 10B tokens also left BLiMP flat. Size does not move it; tokens stop moving it near the cap.
131
- - **Capacity + knowledge drive ARC.** 16M->32M at matched tokens lifted ARC-Easy +3.07pp.
132
- - **The BLiMP ceiling is size-specific, not absolute.** 16M saturates ~70.5 on tokens; a 32M model at 16B tokens reaches BLiMP 73.77 (Path-B v1) and keeps rising - capacity, not data, is the binding constraint at the top.
133
- - **The efficiency sweet-spot is ~32M, not bigger.** Raw scores keep climbing with size (149M: BLiMP 76.99 / ARC 49.66), but efficiency = raw x size-bonus and the bonus falls with size (32M x1.066, 149M x1.000); net, a well-trained 32M outranks a 149M on the board. Beyond ~32M, scale raw-score (data/optimizer/architecture), not parameters.
134
- - **ARC gains are capacity-gated.** ARC-density upweighting was null at 16M (40.87 vs 40.91) but the 32M model reached ARC 44.44; the same data helps only when the model has capacity to exploit it.
135
- - **Efficiency is size-bonus-weighted**, so the smallest model reaching a given raw score ranks highest; ARC is the binding lever toward the top, targeted next via **value residuals** (see below).
136
- - **The ARC lever is value residuals (competitive intel).** The board's #1 model (JugnuLM-110M-R2+) attributes ~+6 ARC-Easy and ~0.18 byte-ppl to value residuals (a ResFormer-style layer-0 value residual) alone β€” the mechanism for the capacity-gated ARC gain. Adopted as the primary ARC lever in the next arch ladder.
137
- - **A fixed small corpus saturates (~16B).** The v1b slope-check (18B on arcmix) confirms diminishing/negative returns from more epochs; the path forward is unique broad-web data (Ultra-FineWeb + DCLM-baseline) plus FineMath, matching the top-3 data stacks.
138
-
139
- ## Roadmap
140
-
141
- - Board entry (**done**): 32M @ #16 (eff 75.51), 16M @ #21 (eff 74.49) β€” maintainer-verified, PR #76 merged. Finding: the arcmix corpus is BLiMP-saturated at ~16B; unique data + architecture are the levers toward the top.
142
- - **ARC lever = value residuals** (ResFormer layer-0 value residual): the #1 model's own card credits ~+6 ARC-Easy to this alone. Primary ARC lever in the next arch ladder (Qwen3 arch: RoPE ΞΈ=100k + RMSNorm + SwiGLU + GQA + QK-Norm + value residuals).
143
- - **Data stack (proven by top-3):** FineWeb-Edu / Ultra-FineWeb (edu backbone) + DCLM-baseline (diverse web) + FineMath-4plus (math). #1 reaches BLiMP 82.52 / ARC 55.13 / byte_ppl 1.8735 with FineWeb-Edu + strong arch (Qwen3 + value residuals + Muon) + a WSD schedule with decay-phase edu upweighting.
144
- - **Escalation:** 2Γ— 64M as a clean A/B (Muon vs AdamW, otherwise identical: winning data stack + full Qwen3+VR arch) β€” two #1 candidates plus a clean optimizer single-factor verdict. Target for #1 at 64M: BLiMP ~82 / ARC ~52 / byte_ppl ~2.0 (eff > 80.2).
145
-
146
- ## Limitations
147
-
148
- Base (not instruction-tuned) research models at 16-32M parameters, English-only. Expect limited factual knowledge and
149
- coherence; not intended for production use.
150
-
151
- ## Provenance
152
-
153
- Full dialectical record, evaluation artifacts and eval-protocol details in labvault
154
- `21_09_GoLLeM-v5-Skalowanie-Glint/` (see `90-Ewaluacja/EvalHarnessParity.md`). Trained on RunPod RTX 5090.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ library_name: pytorch
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - tiny-lm
9
+ - gpt
10
+ - nanogpt
11
+ - glint-tiny-ml-leaderboard
12
+ - english
13
+ datasets:
14
+ - SlayerLab/minimal-en-corpus-5b
15
+ ---
16
+
17
+ # GoLLeM-v5 β€” Tiny English Language Models (16M-64M)
18
+
19
+ Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
20
+ (nanoGPT lineage) trained for the
21
+ [Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
22
+ The repository holds a controlled scaling study (tokens, width, optimizer) and a set of
23
+ **data-attribution experiments** on the 64M model (paired continued-training runs that differ only in data).
24
+
25
+ ## Model details
26
+
27
+ - **Architecture:** 16M/32M = decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings. **64M flagship = Qwen3-style decoder** (RoPE ΞΈ=100k, SwiGLU, RMSNorm, QK-Norm, value residuals).
28
+ - **Sizes:** 16M = 6 layers / d_model 408 / 6 heads (17.4M); 32M = 6 layers / d_model 576 / 9 heads (31.6M); **64M flagship = 14 layers / d_model 576 / 9 heads (62.9M)**.
29
+ - **Context length:** 1024 tokens.
30
+ - **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
31
+ - **Training:** 16M/32M: AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024, seed 1337, bf16 (RTX 5090). 64M flagship: **Muon** (muon-lr 0.02) + AdamW for non-matrix parameters, same cosine schedule, 400k steps.
32
+
33
+ ## Checkpoints
34
+
35
+ | checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
36
+ |---|---|---|---|---|---|---|
37
+ | `bpe16m_3.2B/ckpt.pt` | 17.4M | L6 d408 h6 | 3.2B | 67.40⁺ | 38.22 | 1.2161 |
38
+ | `bpe16m_6B/ckpt.pt` | 17.4M | L6 d408 h6 | 6B | 68.92⁺ | 39.10 | 1.1943 |
39
+ | `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36Β° | 39.52 | 1.1815 |
40
+ | `bpe32m_baseline/ckpt.pt` | 31.6M | L6 d576 h9 | 10B | 70.08Λ’ | 42.59 | 1.124 |
41
+ | `run_16m_expanded/ckpt.pt` (16M board entry) | 17.4M | L6 d408 h6 | 16B† | **70.08** | 40.91 | 1.4193 |
42
+ | `run_32m_16b/ckpt.pt` (32M board entry) | 31.6M | L6 d576 h9 | 16B‑ | **73.48** | **44.44** | 1.3441 |
43
+ | `run_149m/ckpt.pt` (scaling ref) | 149M | β€” | 10BΒ§ | 76.99Β° | 49.66 | 1.2052 |
44
+ | `run_32m_18b/ckpt.pt` (slope-check) | 31.6M | L6 d576 h9 | 18BΒΆ | 72.38Β° | 44.70 | 1.3431 |
45
+ | Muon 32M (results only: `glint_32m_muon_results.json`, checkpoint not published) | 31.6M | L6 d576 h9 | 16Bβ€– | 72.29Β° | 42.89 | 1.3866 |
46
+ | `v1_muon/ckpt_400k.pt` (**64M flagship**) | 62.9M | L14 d576 h9 | 13.1Bβ˜… | **75.83** | **47.94** | 1.246 |
47
+
48
+ **BLiMP harness fix (2026-09-25).** Our evaluation script matched BLiMP configs by substring, so 6 phenomena were counted twice (73,000 pairs instead of 67,000). It is fixed in `glint_parity_eval.py` (exact config match + assertion of 67,000 pairs). The three board entries (bold BLiMP) were recomputed with the fixed script; for each we also reproduced the old 73,000-pair variant, which matches the previously published number to two decimals, so the difference comes from the loader alone (+0.16 to +0.46 pp before the fix). Other rows are **earlier results, not recomputed**:
49
+ - Β° computed with the pre-fix 73,000-pair loader (pair count recorded in the results file); on the three recomputed models the bias was +0.16 to +0.46 pp upward, expect similar here;
50
+ - Λ’ BLiMP on a 4,000-pair random sample, so the error is random (SE ~0.7 pp), not a systematic bias;
51
+ - ⁺ the results file is not archived, so the pair count is unknown; treat as indicative only.
52
+
53
+ † 16M board entry = expanded 8.29B-token corpus (~1.9 epochs). eff **74.33** with the fixed BLiMP (70.08 / 40.91 / byte_ppl 2.6746).
54
+
55
+ ‑ 32M board entry = 32M at 16B tokens on the ARC-MIX corpus. eff **75.41** with the fixed BLiMP (73.48 / 44.44 / byte_ppl 2.5386). It broke the 16M BLiMP ceiling (~70) seen in the token scan: more capacity plus tokens moved both axes.
56
+
57
+ Β§ 149M = scaling reference only (under-trained at 67 tok/param). Highest raw scores in the older runs, but the size multiplier of the efficiency score falls with size, so it ranks lower on eff.
58
+
59
+ ΒΆ Slope-check = 32M at 18B tokens on the same ARC-MIX corpus as the 32M entry. BLiMP 72.38 (lower than at 16B) with byte_ppl flat: more epochs over a fixed corpus did not help. Diagnostic run.
60
+
61
+ β€– 32M with Muon on the expanded corpus. This run changed optimizer and corpus at once, so it cannot isolate Muon. The clean optimizer comparison was later run at 64M (see β˜…). Diagnostic run.
62
+
63
+ β˜… 64M flagship (v1 Muon) = Qwen3-style decoder + value residuals, **Muon** optimizer, ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). Recomputed with the fixed harness: **ARC-Easy 47.94 / BLiMP 75.83 / WikiText-2 byte_ppl 2.372 (BPB 1.246) β†’ eff 75.81.** The maintainer's independent re-benchmark (PR #78, sha256 `59f982c1…` matched) reproduced ARC-Easy 47.94 exactly; its BLiMP 75.99 matches our 73,000-pair variant to two decimals, so it was most likely computed with the same (pre-fix) script. Earlier versions of this card cited eff 77.51 (byte_ppl 2.016, from a wrong bytes-per-token factor 4.755 instead of 3.8605) and later eff ~75.9 (BLiMP 75.99); both are superseded. Muon was at least as good as AdamW in a clean 64M A/B (identical data, architecture and seed; only the optimizer differs); those numbers pre-date the BLiMP fix.
64
+
65
+ ## Data-attribution experiments (64M, 2026-09-25)
66
+
67
+ Question: **does changing the training data in the late phase of training move the efficiency score?** Method: two identical 64M models continue from the same flagship checkpoint (step 320k) with the same recipe, seed and number of steps (80k, to step 400k); **they differ only in the data**. Both are evaluated at 340k / 360k / 380k / 400k with the board protocol and a paired bootstrap over items. Each checkpoint folder has its own README with the full data specification.
68
+
69
+ | run (folder) | data | ARC-E | BLiMP | wiki byte_ppl | eff @400k | note |
70
+ |---|---|---|---|---|---|---|
71
+ | flagship (`v1_muon/`) | ARC-MIX, full run 0β†’400k | 47.94 | 75.83 | 2.3718 | 75.81 | reference |
72
+ | control (`ctrl_arcmix_resume/`) | ARC-MIX | 47.69 | 76.29 | 2.3717 | 75.88 | re-draws the flagship's first 80k steps of data (the data sampler restarted from its seed on resume) |
73
+ | fork-B (`forkB_arcmix_edu/`) | ARC-MIX + educational | 46.34 | 77.20 | 2.3982 | 75.66 | **flawed build**: took only the head of the ARC-MIX file and a different document separator. Not a test of educational data |
74
+ | A (`r3A_arcmix_edu_clean/`) | 45% ARC-MIX + 55% educational (clean build), 2.55B-token pool | 47.31 | 76.22 | 2.3762 | 75.71 | clean pair with B |
75
+ | B (`r3B_arcmix_qa2x/`) | ARC-MIX with Q&A documents Γ—2, 2.62B-token pool | 47.26 | 75.90 | 2.3712 | 75.60 | clean pair with A |
76
+
77
+ - **A βˆ’ B: Ξ”eff +0.11 (95% CI βˆ’0.30 to +0.51)**; at the four checkpoints +0.46 / βˆ’0.13 / βˆ’0.02 / +0.11. No difference on eff. Educational data consistently made WikiText-2 slightly worse (by more than the run-to-run spread); BLiMP was slightly higher but within noise.
78
+ - **Conclusion:** at these data doses and 80k steps of the cosine tail (learning rate 19% β†’ 10% of peak), the effect of the data on eff is **below ~0.5**. None of the runs beats the flagship beyond noise, so none is submitted to the board. This is an upper bound on the effect at this scale of intervention, not evidence that data does not matter.
79
+ - **In progress:** an anchor run (ARC-MIX sample of the same pool size, Q&A Γ—1) and B with a second seed, to measure the Q&A effect cleanly and the run-to-run noise directly.
80
+
81
+ ## Usage
82
+
83
+ These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
84
+ weights. The model class and a board-scoring harness are included in this repo:
85
+
86
+ - `train_gpt_ref.py` β€” GPT definition (rebuild the GPT of the tabled shape, `load_state_dict`, trim logits to vocab 12288).
87
+ - `glint_parity_eval.py` β€” Glint board-scoring forward (256-token clip, raw log-prob) for BLiMP / ARC-Easy / WikiText-2 (BLiMP loader fixed 2026-09-25, see above).
88
+
89
+ ```python
90
+ import torch
91
+ from tokenizers import Tokenizer
92
+ tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
93
+ ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
94
+ state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
95
+ ```
96
+
97
+ **64M checkpoints (`v1_muon/`, and the data-attribution folders) β€” Qwen3-style arch, config embedded in the checkpoint.** Load with the cfg from `ckpt["config"]`, **not** the legacy 5-arg nanoGPT path. `glint_parity_eval.py`'s loader is cfg-aware for both lineages.
98
+
99
+ ```python
100
+ import torch
101
+ from types import SimpleNamespace
102
+ from train_gpt_ref import GPT
103
+ ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
104
+ c = ck["config"] # self-describing: vocab 12288, n_layer 14, n_embd 576, n_head 9, block 1024
105
+ cfg = SimpleNamespace(**c) # Qwen3 flags: rope ΞΈ100k / swiglu / qk_norm / value_residual
106
+ m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
107
+ m.load_state_dict(ck["model"], strict=False) # tied head.weight
108
+ m.eval()
109
+ ```
110
+
111
+ **Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults:
112
+
113
+ ```bash
114
+ # 16M @ expanded corpus
115
+ python train_gpt_ref.py --data-dir <corpus> \
116
+ --n-layer 6 --n-embd 408 --n-head 6 --block 1024 \
117
+ --batch 64 --steps 244141 --lr 6e-4 --min-lr 6e-5 \
118
+ --vocab 12288 --dtype uint16 --seed 1337
119
+ # 32M: --n-embd 576 --n-head 9 (same vocab 12288 / uint16 / BPE tokenizer)
120
+ ```
121
+
122
+ The script's byte-level defaults (`--vocab 256 --dtype uint8`) and its header comment reflect its origin as a **standard-GPT control** compared against an experimental **BDH (fast-weights)** architecture; the leaderboard models here are the standard causal transformer in BPE mode and do **not** use BDH.
123
+
124
+ ## Training data
125
+
126
+ [`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
127
+ β€” ~5.40B BPE-12k tokens, English, decontaminated against the benchmark test sets. A broad high-quality mix:
128
+ FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
129
+ A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) feeds later runs.
130
+
131
+ **ARC-MIX (9.42B).** The 32M board entry and the 64M runs use **ARC-MIX**: a reasoning/knowledge-enriched expansion of the base mixture (9,417,035,832 BPE-12288 tokens) with ARC-relevant science/reasoning/QA web content upweighted (gold ~3Γ—, related ~2Γ—) over the decontaminated base. Same tokenizer, same 13-gram benchmark decontamination (WikiText-2 / BLiMP / ARC test sets).
132
+
133
+ **Earlier FineWeb-Edu-dominant corpus (v2).** A from-scratch 64M run on a FineWeb-Edu-dominant blend scored lower on ARC than ARC-MIX. That build used a different document separator than the rest of our data and a changed recipe (WSD schedule, z-loss, logit cap), so the comparison is confounded and is **not** evidence against educational data trained from scratch.
134
+
135
+ ## Evaluation
136
+
137
+ All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e. the board-comparable definitions:
138
+
139
+ - **BLiMP** β€” 67 configs (train split, 67,000 pairs), each sentence clipped to the first 256 tokens, raw sentence log-prob preference (`good > bad`), no length normalization.
140
+ - **ARC-Easy** β€” test split, zero-shot, raw accuracy over `LL(question + choice) - LL(question)`.
141
+ - **WikiText-2** β€” byte-normalized (byte_ppl / BPB; bytes-per-token 3.8605 for this tokenizer on WikiText-2-raw-v1 test).
142
+
143
+ A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
144
+
145
+ **Board status.** The Glint board still lists the earlier values for our entries (64M: BLiMP 77.84 / byte_ppl 2.016; 32M: BLiMP 73.77; 16M: BLiMP 70.53). The correct values are the ones in this card. With the board's own efficiency formula (`eff = mean(BLiMP, ARC-Easy, normalized-WikiText-2) Γ— size multiplier`, checked line-for-line against the Space source) they give eff **75.81 (64M), 75.41 (32M), 74.33 (16M)**.
146
+
147
+ ## Key findings
148
+
149
+ - **Tokens drive BLiMP at small size, up to a size-specific ceiling.** At 16M, BLiMP rose with tokens (3.2B β†’ 10B) and then flattened near ~70 on the expanded corpus; 32M at 16B tokens moved past that ceiling.
150
+ - **Capacity and knowledge drive ARC.** 16M β†’ 32M at matched tokens lifted ARC-Easy by ~3pp; 64M lifted it further (47.94).
151
+ - **Bigger is not automatically better on eff.** Raw scores keep rising with size, but the efficiency score multiplies by a size bonus that shrinks with parameter count. In our runs 64M is the best eff entry (75.81 vs 75.41 at 32M); a 128M run improves raw quality but the smaller multiplier cancels most of it.
152
+ - **Muon was at least as good as AdamW at 64M** in a clean A/B (same data, architecture, seed; numbers pre-date the BLiMP fix).
153
+ - **Late-phase data changes move eff by less than ~0.5** (see Data-attribution experiments).
154
+ - **Value residuals** are part of the 64M stack. The board's #1 model attributes ~+6 ARC-Easy to them in its own card; we have **not** isolated this effect ourselves.
155
+
156
+ ## Roadmap
157
+
158
+ - Continue training the 64M flagship in segments with a constant-learning-rate phase and a final decay, comparing two data mixes per segment on **validation** sets (never the board test sets) and keeping the winner only when it wins beyond measured run-to-run noise.
159
+ - A clean from-scratch data comparison (same recipe and separator) and synthetic science data for ARC are candidates; both are pending.
160
+
161
+ ## Limitations
162
+
163
+ Base (not instruction-tuned) research models at 16-64M parameters, English-only. Expect limited factual knowledge and
164
+ coherence; not intended for production use.
165
+
166
+ ## Provenance
167
+
168
+ Trained on RunPod RTX 5090. Full evaluation artifacts and protocol details are available on request.
ctrl_arcmix_resume/README.md ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # ctrl_arcmix_resume β€” control arm (ARC-MIX, unchanged data)
2
+
3
+ **Why it exists:** the reference arm of the first data A/B round. It continues the flagship on the same ARC-MIX corpus the flagship was trained on, so that the other arm (`forkB_arcmix_edu/`) can be compared against "more of the same data".
4
+
5
+ **Data:** ARC-MIX 9.42B-token corpus (the flagship's training data), unchanged.
6
+
7
+ **Important caveat:** when this run was resumed from step 320,000, the trainer restarted its batch sampler from the seed. Because the corpus file is identical to the flagship's, this arm re-drew exactly the same training windows the flagship saw in its first 80k steps. It is therefore a *replay* of early data, not fresh data, and comparisons against it carry that confound. A trainer fix (fast-forwarding the sampler on resume) is prepared for future runs.
8
+
9
+ ## Results
10
+
11
+ | checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
12
+ |---|---|---|---|---|---|
13
+ | ckpt_360k.pt | 360000 | 46.63 | 75.76 | 2.3727 | 75.33 |
14
+ | ckpt_400k.pt | 400000 | 47.69 | 76.29 | 2.3717 | 75.88 |
15
+
16
+ ## Common setup
17
+
18
+ - **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs.
19
+ - **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data.
20
+ - **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
21
+ - **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate).
22
+ - **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise.
23
+ - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
forkB_arcmix_edu/README.md ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # forkB_arcmix_edu β€” first data arm (ARC-MIX + educational web text), confounded build
2
+
3
+ **Why it exists:** the treatment arm of the first data A/B round: does adding educational web text (FineWeb-Edu, decontaminated against the evaluation test sets) to ARC-MIX in the last 80k steps improve the model? Mix: about 45% ARC-MIX, 55% educational text, 2.55B-token pool.
4
+
5
+ **Important caveat β€” this is not a clean test of educational data.** The data build had two defects found during training, before evaluation:
6
+ 1. the ARC-MIX half was taken from the beginning of an unshuffled, source-ordered file (about 1.15B tokens) instead of being sampled across the whole corpus, so it contains almost none of the chat/Q&A-format documents present in ARC-MIX;
7
+ 2. educational documents were separated by `<|im_end|>` instead of the corpus end-of-text token.
8
+ The result below is therefore reported only as "this blend vs control". The clean rebuild is `r3A_arcmix_edu_clean/`.
9
+
10
+ ## Results
11
+
12
+ | checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
13
+ |---|---|---|---|---|---|
14
+ | ckpt_360k.pt | 360000 | 45.83 | 76.81 | 2.3996 | 75.35 |
15
+ | ckpt_400k.pt | 400000 | 46.34 | 77.20 | 2.3982 | 75.66 |
16
+
17
+ Versus the control arm (`ctrl_arcmix_resume/`), paired bootstrap on eff: +0.02 [βˆ’0.41, +0.45] at 360k, βˆ’0.22 [βˆ’0.64, +0.21] at 400k β€” no difference. Consistent pattern: BLiMP higher, WikiText-2 byte-perplexity worse. The clean rebuild shows this pattern came from the build defects, not from the educational data.
18
+
19
+ ## Common setup
20
+
21
+ - **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs.
22
+ - **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data.
23
+ - **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
24
+ - **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate).
25
+ - **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise.
26
+ - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
r3A_arcmix_edu_clean/README.md ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # r3A_arcmix_edu_clean β€” arm A of round 3: ARC-MIX + educational text, clean build
2
+
3
+ **Why it exists:** a clean repeat of the first educational-data arm. Same proportion as `forkB_arcmix_edu/` (45.1% ARC-MIX, 54.9% educational text, 2.55B-token pool), with the build fixed: ARC-MIX sampled uniformly across the whole corpus at document boundaries, and every document separated by the corpus end-of-text token. It was trained in parallel with arm B (`r3B_arcmix_qa2x/`) from the same checkpoint.
4
+
5
+ ## Results
6
+
7
+ | checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
8
+ |---|---|---|---|---|---|
9
+ | ckpt_340k.pt | 340000 | 47.35 | 76.13 | 2.3801 | 75.69 |
10
+ | ckpt_360k.pt | 360000 | 47.18 | 76.01 | 2.3839 | 75.58 |
11
+ | ckpt_380k.pt | 380000 | 47.10 | 76.04 | 2.3787 | 75.57 |
12
+ | ckpt_400k.pt | 400000 | 47.31 | 76.22 | 2.3762 | 75.71 |
13
+
14
+ **Verdict (A βˆ’ B, paired bootstrap on eff):** +0.46 / βˆ’0.13 / βˆ’0.02 / +0.11 at 340k / 360k / 380k / 400k; at 400k +0.11 [βˆ’0.30, +0.51]. **No difference on eff.** Across all four checkpoints, arm A has slightly worse WikiText-2 byte-perplexity (a consistent effect of the educational data at this dose) and slightly higher BLiMP (within noise). Conclusion recorded for the study: at this data dose, changing data only in the last 80k steps moves eff by less than about 0.5.
15
+
16
+ ## Common setup
17
+
18
+ - **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs.
19
+ - **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data.
20
+ - **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
21
+ - **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate).
22
+ - **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise.
23
+ - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
r3B_arcmix_qa2x/README.md ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # r3B_arcmix_qa2x β€” arm B of round 3: ARC-MIX with Q&A documents upweighted
2
+
3
+ **Why it exists:** tests whether upweighting question-answering data helps. ARC-MIX sampled uniformly across the whole corpus (2.62B-token pool), with documents in chat/Q&A format (those containing `<|im_start|>`) taken twice, raising their share from 3.7% to 7.1%. No educational data. Trained in parallel with arm A (`r3A_arcmix_edu_clean/`) from the same checkpoint.
4
+
5
+ ## Results
6
+
7
+ | checkpoint | step (read from file) | ARC-Easy | BLiMP | WikiText-2 byte-ppl | eff (board formula, 62.9M) |
8
+ |---|---|---|---|---|---|
9
+ | ckpt_340k.pt | 340000 | 46.42 | 75.71 | 2.3781 | 75.23 |
10
+ | ckpt_360k.pt | 360000 | 47.73 | 75.77 | 2.3744 | 75.71 |
11
+ | ckpt_380k.pt | 380000 | 47.39 | 75.77 | 2.3756 | 75.59 |
12
+ | ckpt_400k.pt | 400000 | 47.26 | 75.90 | 2.3712 | 75.60 |
13
+
14
+ **Verdict:** see `r3A_arcmix_edu_clean/` β€” A βˆ’ B shows no difference on eff (400k: +0.11 [βˆ’0.30, +0.51]). A follow-up round compares this arm against an anchor with the same pool size and no upweighting (`r4K_anchor_arcmix_qa1/`) and against the same data with a different seed (`r4B_arcmix_qa2x_seed1338/`); first readings suggest the 2Γ— Q&A upweight slightly hurts rather than helps, not yet beyond noise.
15
+
16
+ ## Common setup
17
+
18
+ - **Base model:** GoLLeM-v5 64M flagship (`v1_muon/`, 62.9M parameters, 14 layers, d_model 576, 9 heads, RoPE, SwiGLU, RMSNorm, QK-norm, value residual, Muon optimizer). Every arm of the study starts from the flagship checkpoint at step 320,000 and continues to step 400,000 (80k steps, about 2.6B tokens) with the flagship recipe unchanged (same learning-rate schedule, batch, optimizer state and seed). Only the training data differs.
19
+ - **Method:** two arms trained in parallel from the same checkpoint, evaluated at matching steps (360k / 400k for the first round, 340k to 400k for round 3) with the same harness; the difference between arms is attributed to the data.
20
+ - **Evaluation:** `glint_parity_eval.py` in the repository root (fixed version: BLiMP on exactly 67,000 pairs), ARC-Easy test (bare prompt), BLiMP, WikiText-2 test byte-perplexity, eff by the Glint board formula. The step in the tables is read from the checkpoint file, not from its name.
21
+ - **Reference:** flagship `v1_muon/ckpt_400k.pt` scores ARC-Easy 47.94 / BLiMP 75.83 / byte-ppl 2.3718 / eff 75.81. For scale: a second run that differs only in the random seed moved eff by about 0.26 (one seed pair, so a rough indication of run-to-run noise, not a precise estimate).
22
+ - **Status:** research checkpoint, **not a leaderboard submission**. No arm of this study beats the flagship beyond run-to-run noise.
23
+ - **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.