File size: 16,919 Bytes
3ac120d 64a8ba9 3ac120d 64a8ba9 3ac120d 7757b15 179cfda 64a8ba9 3ac120d 179cfda 3ac120d 179cfda ff77f7c 7757b15 6d4bb85 179cfda 64a8ba9 3ac120d b443fdb 02a1315 3ac120d 82b7265 3ac120d 64a8ba9 179cfda 3ac120d ff77f7c 6d4bb85 ff77f7c 179cfda 3ac120d 179cfda 3ac120d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 | ---
license: apache-2.0
language:
- en
library_name: pytorch
pipeline_tag: text-generation
tags:
- tiny-lm
- gpt
- nanogpt
- glint-tiny-ml-leaderboard
- english
datasets:
- SlayerLab/minimal-en-corpus-5b
---
# GoLLeM-v5 β Tiny English Language Models (16M-64M)
Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
(nanoGPT lineage) trained for the
[Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
This repository is a controlled single-factor scaling study: identical architecture/hyper-parameters/seed,
varying only tokens and model width.
## Model details
- **Architecture:** 16M/32M = decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings. **64M flagship = Qwen3-style decoder** (RoPE ΞΈ=100k, SwiGLU, RMSNorm, QK-Norm, value residuals), trained with the **Muon** optimizer.
- **Sizes:** 16M = 6 layers / d_model 408 / 6 heads (17.4M); 32M = 6 layers / d_model 576 / 9 heads (31.4M); **64M flagship = 14 layers / d_model 576 / 9 heads (62.9M)**.
- **Context length:** 1024 tokens.
- **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
- **Training:** AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024 (65,536 tok/step), seed 1337, bf16 (RTX 5090).
## Checkpoints
| checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
|---|---|---|---|---|---|---|
| `bpe16m_3.2B/ckpt.pt` | 17.4M | L6 d408 h6 | 3.2B | 67.40 | 38.22 | 1.2161 |
| `bpe16m_6B/ckpt.pt` | 17.4M | L6 d408 h6 | 6B | 68.92 | 39.10 | 1.1943 |
| `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 |
| `bpe32m_baseline/ckpt.pt` | 31.4M | L6 d576 h9 | 10B | 70.08 | 42.59 | 1.124 |
| `run_16m_expanded/ckpt.pt` (crown) | 17.4M | L6 d408 h6 | 16Bβ | 70.53 | 40.91 | 1.4193 |
| `run_32m_16b/ckpt.pt` (Path-B v1) | 31.4M | L6 d576 h9 | 16Bβ‘ | **73.77** | **44.44** | 1.3441 |
| `run_149m/ckpt.pt` (scaling ref) | 149M | β | 10BΒ§ | 76.99 | 49.66 | 1.2052 |
| `run_32m_18b/ckpt.pt` (v1b slope-check) | 31.6M | L6 d576 h9 | 18BΒΆ | 72.38 | 44.70 | 1.3431 |
| `run_32m_muon/ckpt.pt` (Muon optimizer) | 31.6M | L6 d576 h9 | 16Bβ | 72.29 | 42.89 | 1.3866 |
| `v1_muon/ckpt_400k.pt` (**#6 flagship** β Qwen3+Muon+VR) | 62.9M | L14 d576 h9 | 13.1Bβ
| **77.84** | **47.94** | 1.012 |
β crown = expanded 8.29B-token corpus (~1.9 epochs). This is the published **16M board entry: confirmed #21** (eff 74.49 β 70.53 / 40.91 / byte_ppl 2.6746). Earlier recon estimated #20 (an optimistic #14 used a wrong wiki estimate before the exact byte_ppl correction); the official merge landed #21.
β‘ Path-B v1 = 32M at 16B tokens (ARC-MIX corpus): **breaks the 16M BLiMP ceiling** (70.5 -> 73.77) and lifts ARC to 44.44. This is the published **32M board entry: confirmed #16** (eff 75.51); recon estimated #15, official merge landed #16. The 70.5 cap was 16M-specific, not absolute; more capacity + tokens moves both axes.
Β§ 149M = scaling reference only (heavily under-trained at 67 tok/param). Highest raw scores (BLiMP 76.99 / ARC 49.66) but board-recon **#20/74** β the efficiency size-bonus caps at ~32M, so bigger models score higher raw but rank lower on eff. The eff sweet-spot is ~32M; the lever toward the top is raw-score at 32M (architecture / optimizer / tokens), not more size.
ΒΆ v1b slope-check = 32M at 18B tokens on the **same arcmix corpus** as Path-B v1. BLiMP 72.38 (β1.39 vs v1@16B) with byte_ppl flat (2.537 vs 2.539) β **the arcmix corpus is saturated at ~16B**: more epochs (~1.9) over-cycle and mildly hurt BLiMP. This pre-registered slope-check refutes "train longer" on a fixed corpus; the next gain needs **unique** data (broad web), not re-cycled epochs. Diagnostic run (ckpt on request).
β Muon optimizer = 32M at 16B tokens, Muon optimizer (muon-lr 0.02), on the expanded 8.29B corpus. BLiMP 72.29 / ARC 42.89 / byte_ppl 2.615. **Optimizer verdict = inconclusive (confounded design):** this run also changed corpus (expanded vs Path-B's arcmix), so the β1.48 BLiMP mixes optimizer *and* data and cannot isolate Muon. A clean Muon single-factor is deferred to the 64M A/B (Muon vs AdamW, corpus held fixed). Diagnostic run.
β
64M flagship (v1 Muon) = 62.9M, **Qwen3-style decoder** (RoPE ΞΈ=100k + SwiGLU + RMSNorm + QK-Norm + value residuals), **Muon** optimizer (muon-lr 0.02, cosine), ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). **Board entry: #6 GoLLeM-v5 64M** (eff 77.51 β BLiMP 77.84 / ARC-Easy 47.94 / WikiText-2 byte_ppl 2.016, BPB 1.012). Recompute-verified **board-protocol** (Glint bare-prompt ARC β `LL(choice|q)` argmax, full BLiMP-67k, wiki byte_ppl, no-BOS β matching the maintainer's `glint_parity` harness), ckpt sha256 `59f982c1β¦`. A clean single-factor scale-up of the 32M Path-B recipe (same BPE-12288 tokenizer / arcmix data lineage; only size + the Qwen3+Muon+VR arch differ) that lifts eff 75.51 β 77.51. Also settles the clean **Muon verdict**: at 64M with corpus held fixed (Muon vs AdamW A/B), Muon wins on byte_ppl + BLiMP.
## Usage
These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
weights. The model class and a ready board-scoring harness are included in this repo:
- `train_gpt_ref.py` β GPT definition (rebuild the GPT of the tabled shape, `load_state_dict`, trim logits to vocab 12288).
- `glint_parity_eval.py` β the exact Glint board-scoring forward (256-token clip, raw log-prob) for BLiMP / ARC-Easy / WikiText-2.
```python
import torch
from tokenizers import Tokenizer
tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
```
**64M flagship (Qwen3-arch) β load from the checkpoint's embedded config.** The 64M is a **Qwen3-style decoder** (RoPE + SwiGLU + RMSNorm + QK-Norm + value residuals), *not* the nanoGPT of the 16M/32M. Its architecture is self-described in `ckpt["config"]`; rebuild from that (do **not** load it as the plain 16M/32M GPT β a plain-GPT loader mis-loads to silent wrong logits):
```python
import torch
from types import SimpleNamespace
from train_gpt_ref import GPT
ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
c = ck["config"] # rope/swiglu/rmsnorm/qk_norm/value_residual; n_head 9, L14, d576, block 1024
cfg = SimpleNamespace(**c)
m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
m.load_state_dict(ck["model"], strict=False) # tied head
m.eval()
```
The bundled `glint_parity_eval.py` is arch-aware (auto-detects `ckpt["config"]`) β run it directly on `v1_muon/ckpt_400k.pt` to reproduce the board numbers (BLiMP 77.84 / ARC-Easy 47.94 bare-prompt / eff 77.51).
**64M flagship (`v1_muon/ckpt_400k.pt`) β Qwen3-style arch, config embedded in the checkpoint.** Load with the cfg from `ckpt["config"]` (RoPE / SwiGLU / QK-Norm / value-residual, `n_head=9` / L14 / d576 / block 1024 are self-describing), **not** the legacy 5-arg nanoGPT path. `glint_parity_eval.py`'s loader is cfg-aware for both lineages.
```python
import torch
from types import SimpleNamespace
from train_gpt_ref import GPT
ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
c = ck["config"] # self-describing: vocab 12288, n_layer 14, n_embd 576, n_head 9, block 1024
cfg = SimpleNamespace(**c) # Qwen3 flags: rope ΞΈ100k / swiglu / qk_norm / value_residual
m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
m.load_state_dict(ck["model"], strict=False) # tied head.weight
m.eval()
```
**Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above):
```bash
# crown 16M @ expanded corpus
python train_gpt_ref.py --data-dir <corpus> \
--n-layer 6 --n-embd 408 --n-head 6 --block 1024 \
--batch 64 --steps 244141 --lr 6e-4 --min-lr 6e-5 \
--vocab 12288 --dtype uint16 --seed 1337
# Path-B 32M: --n-embd 576 --n-head 9 (same vocab 12288 / uint16 / BPE tokenizer)
```
The `--vocab 12288 --dtype uint16` flags select BPE-12k over uint16 token bins. The script's byte-level defaults (`--vocab 256 --dtype uint8`) and its header comment reflect its origin as a **standard-GPT control** compared against an experimental **BDH (fast-weights)** architecture β the leaderboard models here are the standard causal transformer in BPE mode and do **not** use BDH.
## Training data
[`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
β ~5.40B BPE-12k tokens, English, decontaminated against the benchmark test sets. A broad high-quality mix:
FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) feeds later runs.
**ARC-MIX (9.42B).** The 32M Path-B and 64M A/B runs use **ARC-MIX** β a reasoning/knowledge-enriched expansion of the base mixture (9,417,035,832 BPE-12288 tokens): ARC-relevant science/reasoning/QA web content upweighted (gold ~3Γ, related ~2Γ) over the decontaminated base, to push the ARC-Easy axis (the binding efficiency constraint at this scale; capacity-gated per finding W11). Same BPE-12288 tokenizer, same 13-gram benchmark decontamination (WikiText-2 / BLiMP / ARC). The v2 #1-shot corpus moves to a FineWeb-Edu-dominant blend (β₯60% FineWeb-Edu + DCLM-baseline + FineMath-4plus, ~20B unique), per the 8M data-screen (FineWeb-Edu won BLiMP) and top-3 competitor recipes.
**Training budget and epochs.** 16M trained on 16B tokens *seen* is intentional, not a chart error. The leaderboard scores efficiency = quality at a fixed tiny size, so you over-train to squeeze max quality from frozen capacity. Chinchilla-optimal (about 20x params = 0.32B for 16M) minimizes compute-optimal loss, which is NOT the leaderboard objective; top models train many tokens-per-param too. 16B seen over 8.29B unique corpus = about 1.9 epochs (each token seen ~1.9x, under the 2x repeat-degradation limit; val-loss healthy, zero memorization). Note: the crown 16B point changed BOTH tokens and corpus (5.4B to 8.29B expanded), so on the token-scan chart it is marked separately (star + dashed) and is not a pure token step. BLiMP saturation holds regardless: crown BLiMP 70.53 is below even the token-only projection 71.6.
## Evaluation
All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e. the board-comparable definitions:
- **BLiMP** β 67 configs (train split), each sentence clipped to the first 256 tokens, raw sentence log-prob preference (`good > bad`), no length normalization.
- **ARC-Easy** β test split, zero-shot, raw accuracy over `LL(question + choice) - LL(question)`.
- **WikiText-2** β byte-normalized bits-per-byte (the board's `wiki` field is byte-scale; token-perplexity is tokenizer-dependent and not directly comparable across models).
A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
**#6 β GoLLeM-v5 64M flagship (eff 77.51), board-protocol-verified (Glint bare-prompt ARC, 2026-09-24).** The 64M Muon model (Qwen3 arch + value residuals, ARC-MIX 9.42B) lands **#6 on the Glint Tiny-ML Leaderboard** β BLiMP 77.84 / ARC-Easy 47.94 / WikiText-2 byte_ppl 2.016 β behind only four 90β143M models and Glint-1.3 (982K, #5; razor-thin, eff 77.58 vs 77.51). It is the **strongest dense 64M entry** on the board, a **#16 β #6 jump** from the 32M. A clean single-factor scale-up (32M β 64M, same data/tokenizer lineage) plus the Qwen3+Muon+value-residual stack lifted eff 75.51 β 77.51. Numbers are recompute-verified **board-native** (bare-prompt ARC, matching the maintainer's `glint_parity` harness β **not** lm-eval), reproducible from the checkpoint.
**Positioning β CONFIRMED, on the board.** PR #76 was merged into the Glint Tiny-ML Leaderboard (2026-09-23), maintainer-verified (checkpoints loaded directly; params confirmed: 32M = 31,601,664, 16M = 17,449,344 deduped tied-embeddings; architecture matches `train_gpt_ref.py`, standard nanoGPT BPE-12288). Official standings: **#16 GoLLeM-v5 32M** (eff 75.51 β BLiMP 73.77 / ARC-Easy 44.44 / WikiText-2 byte_ppl 2.5386, 16B tok) and **#21 GoLLeM-v5 16M** (eff 74.49 β BLiMP 70.53 / ARC 40.91 / byte_ppl 2.6746). The board efficiency formula was reverse-engineered and then confirmed **line-for-line against the Space source** (reproduces the displayed eff exactly, 3/3 checked models to 2 decimals): `eff = mean(BLiMP, ARC-Easy, normalized-WikiText-2) Γ size-bonus`, where the size-bonus runs 1.0Γ (largest on board) to 1.5Γ (smallest) on a log-parameter scale. Our 32M carries a 1.065Γ size-bonus vs the #1's 1.013Γ β a ~5% efficiency edge at equal raw metrics.
## Key findings (single-factor study)
- **Tokens drive BLiMP, not size β up to a ceiling.** BLiMP climbs with tokens (~+1.8pp per doubling, 3.2B->10B) then **saturates at 16M's ~70.5 ceiling** (crown 16B: 70.53, +0.17 over 10B β flat, below the token-only projection 71.6); 16M->32M at matched 10B tokens also left BLiMP flat. Size does not move it; tokens stop moving it near the cap.
- **Capacity + knowledge drive ARC.** 16M->32M at matched tokens lifted ARC-Easy +3.07pp.
- **The BLiMP ceiling is size-specific, not absolute.** 16M saturates ~70.5 on tokens; a 32M model at 16B tokens reaches BLiMP 73.77 (Path-B v1) and keeps rising - capacity, not data, is the binding constraint at the top.
- **The efficiency sweet-spot is ~32M, not bigger.** Raw scores keep climbing with size (149M: BLiMP 76.99 / ARC 49.66), but efficiency = raw x size-bonus and the bonus falls with size (32M x1.066, 149M x1.000); net, a well-trained 32M outranks a 149M on the board. Beyond ~32M, scale raw-score (data/optimizer/architecture), not parameters.
- **ARC gains are capacity-gated.** ARC-density upweighting was null at 16M (40.87 vs 40.91) but the 32M model reached ARC 44.44; the same data helps only when the model has capacity to exploit it.
- **Efficiency is size-bonus-weighted**, so the smallest model reaching a given raw score ranks highest; ARC is the binding lever toward the top, targeted next via **value residuals** (see below).
- **The ARC lever is value residuals (competitive intel).** The board's #1 model (JugnuLM-110M-R2+) attributes ~+6 ARC-Easy and ~0.18 byte-ppl to value residuals (a ResFormer-style layer-0 value residual) alone β the mechanism for the capacity-gated ARC gain. Adopted as the primary ARC lever in the next arch ladder.
- **A fixed small corpus saturates (~16B).** The v1b slope-check (18B on arcmix) confirms diminishing/negative returns from more epochs; the path forward is unique broad-web data (Ultra-FineWeb + DCLM-baseline) plus FineMath, matching the top-3 data stacks.
## Roadmap
- Board entry (**done**): 32M @ #16 (eff 75.51), 16M @ #21 (eff 74.49) β maintainer-verified, PR #76 merged. Finding: the arcmix corpus is BLiMP-saturated at ~16B; unique data + architecture are the levers toward the top.
- **ARC lever = value residuals** (ResFormer layer-0 value residual): the #1 model's own card credits ~+6 ARC-Easy to this alone. Primary ARC lever in the next arch ladder (Qwen3 arch: RoPE ΞΈ=100k + RMSNorm + SwiGLU + GQA + QK-Norm + value residuals).
- **Data stack (proven by top-3):** FineWeb-Edu / Ultra-FineWeb (edu backbone) + DCLM-baseline (diverse web) + FineMath-4plus (math). #1 reaches BLiMP 82.52 / ARC 55.13 / byte_ppl 1.8735 with FineWeb-Edu + strong arch (Qwen3 + value residuals + Muon) + a WSD schedule with decay-phase edu upweighting.
- **Escalation:** 2Γ 64M as a clean A/B (Muon vs AdamW, otherwise identical: winning data stack + full Qwen3+VR arch) β two #1 candidates plus a clean optimizer single-factor verdict. Target for #1 at 64M: BLiMP ~82 / ARC ~52 / byte_ppl ~2.0 (eff > 80.2).
## Limitations
Base (not instruction-tuned) research models at 16-32M parameters, English-only. Expect limited factual knowledge and
coherence; not intended for production use.
## Provenance
Full dialectical record, evaluation artifacts and eval-protocol details in labvault
`21_09_GoLLeM-v5-Skalowanie-Glint/` (see `90-Ewaluacja/EvalHarnessParity.md`). Trained on RunPod RTX 5090.
|