gollem-v5-ckpts / README.md
Maggio33's picture
card: add 64M Qwen3 load-recipe (maintainer audit load-safety)
b443fdb verified
|
Raw History Blame
16.9 kB
---
license: apache-2.0
language:
- en
library_name: pytorch
pipeline_tag: text-generation
tags:
- tiny-lm
- gpt
- nanogpt
- glint-tiny-ml-leaderboard
- english
datasets:
- SlayerLab/minimal-en-corpus-5b
---
# GoLLeM-v5 β€” Tiny English Language Models (16M-64M)
Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
(nanoGPT lineage) trained for the
[Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
This repository is a controlled single-factor scaling study: identical architecture/hyper-parameters/seed,
varying only tokens and model width.
## Model details
- **Architecture:** 16M/32M = decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings. **64M flagship = Qwen3-style decoder** (RoPE ΞΈ=100k, SwiGLU, RMSNorm, QK-Norm, value residuals), trained with the **Muon** optimizer.
- **Sizes:** 16M = 6 layers / d_model 408 / 6 heads (17.4M); 32M = 6 layers / d_model 576 / 9 heads (31.4M); **64M flagship = 14 layers / d_model 576 / 9 heads (62.9M)**.
- **Context length:** 1024 tokens.
- **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
- **Training:** AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024 (65,536 tok/step), seed 1337, bf16 (RTX 5090).
## Checkpoints
| checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
|---|---|---|---|---|---|---|
| `bpe16m_3.2B/ckpt.pt` | 17.4M | L6 d408 h6 | 3.2B | 67.40 | 38.22 | 1.2161 |
| `bpe16m_6B/ckpt.pt` | 17.4M | L6 d408 h6 | 6B | 68.92 | 39.10 | 1.1943 |
| `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 |
| `bpe32m_baseline/ckpt.pt` | 31.4M | L6 d576 h9 | 10B | 70.08 | 42.59 | 1.124 |
| `run_16m_expanded/ckpt.pt` (crown) | 17.4M | L6 d408 h6 | 16B† | 70.53 | 40.91 | 1.4193 |
| `run_32m_16b/ckpt.pt` (Path-B v1) | 31.4M | L6 d576 h9 | 16B‑ | **73.77** | **44.44** | 1.3441 |
| `run_149m/ckpt.pt` (scaling ref) | 149M | β€” | 10BΒ§ | 76.99 | 49.66 | 1.2052 |
| `run_32m_18b/ckpt.pt` (v1b slope-check) | 31.6M | L6 d576 h9 | 18BΒΆ | 72.38 | 44.70 | 1.3431 |
| `run_32m_muon/ckpt.pt` (Muon optimizer) | 31.6M | L6 d576 h9 | 16Bβ€– | 72.29 | 42.89 | 1.3866 |
| `v1_muon/ckpt_400k.pt` (**#6 flagship** β€” Qwen3+Muon+VR) | 62.9M | L14 d576 h9 | 13.1Bβ˜… | **77.84** | **47.94** | 1.012 |
† crown = expanded 8.29B-token corpus (~1.9 epochs). This is the published **16M board entry: confirmed #21** (eff 74.49 β€” 70.53 / 40.91 / byte_ppl 2.6746). Earlier recon estimated #20 (an optimistic #14 used a wrong wiki estimate before the exact byte_ppl correction); the official merge landed #21.
‑ Path-B v1 = 32M at 16B tokens (ARC-MIX corpus): **breaks the 16M BLiMP ceiling** (70.5 -> 73.77) and lifts ARC to 44.44. This is the published **32M board entry: confirmed #16** (eff 75.51); recon estimated #15, official merge landed #16. The 70.5 cap was 16M-specific, not absolute; more capacity + tokens moves both axes.
Β§ 149M = scaling reference only (heavily under-trained at 67 tok/param). Highest raw scores (BLiMP 76.99 / ARC 49.66) but board-recon **#20/74** β€” the efficiency size-bonus caps at ~32M, so bigger models score higher raw but rank lower on eff. The eff sweet-spot is ~32M; the lever toward the top is raw-score at 32M (architecture / optimizer / tokens), not more size.
ΒΆ v1b slope-check = 32M at 18B tokens on the **same arcmix corpus** as Path-B v1. BLiMP 72.38 (βˆ’1.39 vs v1@16B) with byte_ppl flat (2.537 vs 2.539) β€” **the arcmix corpus is saturated at ~16B**: more epochs (~1.9) over-cycle and mildly hurt BLiMP. This pre-registered slope-check refutes "train longer" on a fixed corpus; the next gain needs **unique** data (broad web), not re-cycled epochs. Diagnostic run (ckpt on request).
β€– Muon optimizer = 32M at 16B tokens, Muon optimizer (muon-lr 0.02), on the expanded 8.29B corpus. BLiMP 72.29 / ARC 42.89 / byte_ppl 2.615. **Optimizer verdict = inconclusive (confounded design):** this run also changed corpus (expanded vs Path-B's arcmix), so the βˆ’1.48 BLiMP mixes optimizer *and* data and cannot isolate Muon. A clean Muon single-factor is deferred to the 64M A/B (Muon vs AdamW, corpus held fixed). Diagnostic run.
β˜… 64M flagship (v1 Muon) = 62.9M, **Qwen3-style decoder** (RoPE ΞΈ=100k + SwiGLU + RMSNorm + QK-Norm + value residuals), **Muon** optimizer (muon-lr 0.02, cosine), ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). **Board entry: #6 GoLLeM-v5 64M** (eff 77.51 β€” BLiMP 77.84 / ARC-Easy 47.94 / WikiText-2 byte_ppl 2.016, BPB 1.012). Recompute-verified **board-protocol** (Glint bare-prompt ARC β€” `LL(choice|q)` argmax, full BLiMP-67k, wiki byte_ppl, no-BOS β€” matching the maintainer's `glint_parity` harness), ckpt sha256 `59f982c1…`. A clean single-factor scale-up of the 32M Path-B recipe (same BPE-12288 tokenizer / arcmix data lineage; only size + the Qwen3+Muon+VR arch differ) that lifts eff 75.51 β†’ 77.51. Also settles the clean **Muon verdict**: at 64M with corpus held fixed (Muon vs AdamW A/B), Muon wins on byte_ppl + BLiMP.
## Usage
These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
weights. The model class and a ready board-scoring harness are included in this repo:
- `train_gpt_ref.py` β€” GPT definition (rebuild the GPT of the tabled shape, `load_state_dict`, trim logits to vocab 12288).
- `glint_parity_eval.py` β€” the exact Glint board-scoring forward (256-token clip, raw log-prob) for BLiMP / ARC-Easy / WikiText-2.
```python
import torch
from tokenizers import Tokenizer
tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
```
**64M flagship (Qwen3-arch) β€” load from the checkpoint's embedded config.** The 64M is a **Qwen3-style decoder** (RoPE + SwiGLU + RMSNorm + QK-Norm + value residuals), *not* the nanoGPT of the 16M/32M. Its architecture is self-described in `ckpt["config"]`; rebuild from that (do **not** load it as the plain 16M/32M GPT β€” a plain-GPT loader mis-loads to silent wrong logits):
```python
import torch
from types import SimpleNamespace
from train_gpt_ref import GPT
ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
c = ck["config"] # rope/swiglu/rmsnorm/qk_norm/value_residual; n_head 9, L14, d576, block 1024
cfg = SimpleNamespace(**c)
m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
m.load_state_dict(ck["model"], strict=False) # tied head
m.eval()
```
The bundled `glint_parity_eval.py` is arch-aware (auto-detects `ckpt["config"]`) β€” run it directly on `v1_muon/ckpt_400k.pt` to reproduce the board numbers (BLiMP 77.84 / ARC-Easy 47.94 bare-prompt / eff 77.51).
**64M flagship (`v1_muon/ckpt_400k.pt`) β€” Qwen3-style arch, config embedded in the checkpoint.** Load with the cfg from `ckpt["config"]` (RoPE / SwiGLU / QK-Norm / value-residual, `n_head=9` / L14 / d576 / block 1024 are self-describing), **not** the legacy 5-arg nanoGPT path. `glint_parity_eval.py`'s loader is cfg-aware for both lineages.
```python
import torch
from types import SimpleNamespace
from train_gpt_ref import GPT
ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
c = ck["config"] # self-describing: vocab 12288, n_layer 14, n_embd 576, n_head 9, block 1024
cfg = SimpleNamespace(**c) # Qwen3 flags: rope ΞΈ100k / swiglu / qk_norm / value_residual
m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
m.load_state_dict(ck["model"], strict=False) # tied head.weight
m.eval()
```
**Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above):
```bash
# crown 16M @ expanded corpus
python train_gpt_ref.py --data-dir <corpus> \
--n-layer 6 --n-embd 408 --n-head 6 --block 1024 \
--batch 64 --steps 244141 --lr 6e-4 --min-lr 6e-5 \
--vocab 12288 --dtype uint16 --seed 1337
# Path-B 32M: --n-embd 576 --n-head 9 (same vocab 12288 / uint16 / BPE tokenizer)
```
The `--vocab 12288 --dtype uint16` flags select BPE-12k over uint16 token bins. The script's byte-level defaults (`--vocab 256 --dtype uint8`) and its header comment reflect its origin as a **standard-GPT control** compared against an experimental **BDH (fast-weights)** architecture β€” the leaderboard models here are the standard causal transformer in BPE mode and do **not** use BDH.
## Training data
[`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
β€” ~5.40B BPE-12k tokens, English, decontaminated against the benchmark test sets. A broad high-quality mix:
FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) feeds later runs.
**ARC-MIX (9.42B).** The 32M Path-B and 64M A/B runs use **ARC-MIX** β€” a reasoning/knowledge-enriched expansion of the base mixture (9,417,035,832 BPE-12288 tokens): ARC-relevant science/reasoning/QA web content upweighted (gold ~3Γ—, related ~2Γ—) over the decontaminated base, to push the ARC-Easy axis (the binding efficiency constraint at this scale; capacity-gated per finding W11). Same BPE-12288 tokenizer, same 13-gram benchmark decontamination (WikiText-2 / BLiMP / ARC). The v2 #1-shot corpus moves to a FineWeb-Edu-dominant blend (β‰₯60% FineWeb-Edu + DCLM-baseline + FineMath-4plus, ~20B unique), per the 8M data-screen (FineWeb-Edu won BLiMP) and top-3 competitor recipes.
**Training budget and epochs.** 16M trained on 16B tokens *seen* is intentional, not a chart error. The leaderboard scores efficiency = quality at a fixed tiny size, so you over-train to squeeze max quality from frozen capacity. Chinchilla-optimal (about 20x params = 0.32B for 16M) minimizes compute-optimal loss, which is NOT the leaderboard objective; top models train many tokens-per-param too. 16B seen over 8.29B unique corpus = about 1.9 epochs (each token seen ~1.9x, under the 2x repeat-degradation limit; val-loss healthy, zero memorization). Note: the crown 16B point changed BOTH tokens and corpus (5.4B to 8.29B expanded), so on the token-scan chart it is marked separately (star + dashed) and is not a pure token step. BLiMP saturation holds regardless: crown BLiMP 70.53 is below even the token-only projection 71.6.
## Evaluation
All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e. the board-comparable definitions:
- **BLiMP** β€” 67 configs (train split), each sentence clipped to the first 256 tokens, raw sentence log-prob preference (`good > bad`), no length normalization.
- **ARC-Easy** β€” test split, zero-shot, raw accuracy over `LL(question + choice) - LL(question)`.
- **WikiText-2** β€” byte-normalized bits-per-byte (the board's `wiki` field is byte-scale; token-perplexity is tokenizer-dependent and not directly comparable across models).
A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
**#6 β€” GoLLeM-v5 64M flagship (eff 77.51), board-protocol-verified (Glint bare-prompt ARC, 2026-09-24).** The 64M Muon model (Qwen3 arch + value residuals, ARC-MIX 9.42B) lands **#6 on the Glint Tiny-ML Leaderboard** β€” BLiMP 77.84 / ARC-Easy 47.94 / WikiText-2 byte_ppl 2.016 β€” behind only four 90–143M models and Glint-1.3 (982K, #5; razor-thin, eff 77.58 vs 77.51). It is the **strongest dense 64M entry** on the board, a **#16 β†’ #6 jump** from the 32M. A clean single-factor scale-up (32M β†’ 64M, same data/tokenizer lineage) plus the Qwen3+Muon+value-residual stack lifted eff 75.51 β†’ 77.51. Numbers are recompute-verified **board-native** (bare-prompt ARC, matching the maintainer's `glint_parity` harness β€” **not** lm-eval), reproducible from the checkpoint.
**Positioning β€” CONFIRMED, on the board.** PR #76 was merged into the Glint Tiny-ML Leaderboard (2026-09-23), maintainer-verified (checkpoints loaded directly; params confirmed: 32M = 31,601,664, 16M = 17,449,344 deduped tied-embeddings; architecture matches `train_gpt_ref.py`, standard nanoGPT BPE-12288). Official standings: **#16 GoLLeM-v5 32M** (eff 75.51 β€” BLiMP 73.77 / ARC-Easy 44.44 / WikiText-2 byte_ppl 2.5386, 16B tok) and **#21 GoLLeM-v5 16M** (eff 74.49 β€” BLiMP 70.53 / ARC 40.91 / byte_ppl 2.6746). The board efficiency formula was reverse-engineered and then confirmed **line-for-line against the Space source** (reproduces the displayed eff exactly, 3/3 checked models to 2 decimals): `eff = mean(BLiMP, ARC-Easy, normalized-WikiText-2) Γ— size-bonus`, where the size-bonus runs 1.0Γ— (largest on board) to 1.5Γ— (smallest) on a log-parameter scale. Our 32M carries a 1.065Γ— size-bonus vs the #1's 1.013Γ— β€” a ~5% efficiency edge at equal raw metrics.
## Key findings (single-factor study)
- **Tokens drive BLiMP, not size β€” up to a ceiling.** BLiMP climbs with tokens (~+1.8pp per doubling, 3.2B->10B) then **saturates at 16M's ~70.5 ceiling** (crown 16B: 70.53, +0.17 over 10B β€” flat, below the token-only projection 71.6); 16M->32M at matched 10B tokens also left BLiMP flat. Size does not move it; tokens stop moving it near the cap.
- **Capacity + knowledge drive ARC.** 16M->32M at matched tokens lifted ARC-Easy +3.07pp.
- **The BLiMP ceiling is size-specific, not absolute.** 16M saturates ~70.5 on tokens; a 32M model at 16B tokens reaches BLiMP 73.77 (Path-B v1) and keeps rising - capacity, not data, is the binding constraint at the top.
- **The efficiency sweet-spot is ~32M, not bigger.** Raw scores keep climbing with size (149M: BLiMP 76.99 / ARC 49.66), but efficiency = raw x size-bonus and the bonus falls with size (32M x1.066, 149M x1.000); net, a well-trained 32M outranks a 149M on the board. Beyond ~32M, scale raw-score (data/optimizer/architecture), not parameters.
- **ARC gains are capacity-gated.** ARC-density upweighting was null at 16M (40.87 vs 40.91) but the 32M model reached ARC 44.44; the same data helps only when the model has capacity to exploit it.
- **Efficiency is size-bonus-weighted**, so the smallest model reaching a given raw score ranks highest; ARC is the binding lever toward the top, targeted next via **value residuals** (see below).
- **The ARC lever is value residuals (competitive intel).** The board's #1 model (JugnuLM-110M-R2+) attributes ~+6 ARC-Easy and ~0.18 byte-ppl to value residuals (a ResFormer-style layer-0 value residual) alone β€” the mechanism for the capacity-gated ARC gain. Adopted as the primary ARC lever in the next arch ladder.
- **A fixed small corpus saturates (~16B).** The v1b slope-check (18B on arcmix) confirms diminishing/negative returns from more epochs; the path forward is unique broad-web data (Ultra-FineWeb + DCLM-baseline) plus FineMath, matching the top-3 data stacks.
## Roadmap
- Board entry (**done**): 32M @ #16 (eff 75.51), 16M @ #21 (eff 74.49) β€” maintainer-verified, PR #76 merged. Finding: the arcmix corpus is BLiMP-saturated at ~16B; unique data + architecture are the levers toward the top.
- **ARC lever = value residuals** (ResFormer layer-0 value residual): the #1 model's own card credits ~+6 ARC-Easy to this alone. Primary ARC lever in the next arch ladder (Qwen3 arch: RoPE ΞΈ=100k + RMSNorm + SwiGLU + GQA + QK-Norm + value residuals).
- **Data stack (proven by top-3):** FineWeb-Edu / Ultra-FineWeb (edu backbone) + DCLM-baseline (diverse web) + FineMath-4plus (math). #1 reaches BLiMP 82.52 / ARC 55.13 / byte_ppl 1.8735 with FineWeb-Edu + strong arch (Qwen3 + value residuals + Muon) + a WSD schedule with decay-phase edu upweighting.
- **Escalation:** 2Γ— 64M as a clean A/B (Muon vs AdamW, otherwise identical: winning data stack + full Qwen3+VR arch) β€” two #1 candidates plus a clean optimizer single-factor verdict. Target for #1 at 64M: BLiMP ~82 / ARC ~52 / byte_ppl ~2.0 (eff > 80.2).
## Limitations
Base (not instruction-tuned) research models at 16-32M parameters, English-only. Expect limited factual knowledge and
coherence; not intended for production use.
## Provenance
Full dialectical record, evaluation artifacts and eval-protocol details in labvault
`21_09_GoLLeM-v5-Skalowanie-Glint/` (see `90-Ewaluacja/EvalHarnessParity.md`). Trained on RunPod RTX 5090.