--- license: apache-2.0 language: - en library_name: pytorch pipeline_tag: text-generation tags: - tiny-lm - gpt - nanogpt - glint-tiny-ml-leaderboard - english datasets: - SlayerLab/minimal-en-corpus-5b --- # GoLLeM-v5 — Tiny English Language Models (16M-32M) Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders (nanoGPT lineage) trained for the [Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard). This repository is a controlled single-factor scaling study: identical architecture/hyper-parameters/seed, varying only tokens and model width. ## Model details - **Architecture:** decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings. - **Sizes:** 16M variant = 6 layers / d_model 408 / 6 heads (17.4M params); 32M variant = 6 layers / d_model 576 / 9 heads (31.4M params). - **Context length:** 1024 tokens. - **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints. - **Training:** AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024 (65,536 tok/step), seed 1337, bf16 (RTX 5090). ## Checkpoints | checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB | |---|---|---|---|---|---|---| | `bpe16m_3.2B/ckpt.pt` | 17.4M | L6 d408 h6 | 3.2B | 67.40 | 38.22 | 1.2161 | | `bpe16m_6B/ckpt.pt` | 17.4M | L6 d408 h6 | 6B | 68.92 | 39.10 | 1.1943 | | `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 | | `bpe32m_baseline/ckpt.pt` | 31.4M | L6 d576 h9 | 10B | 70.08 | 42.59 | 1.124 | | `run_16m_expanded/ckpt.pt` (crown) | 17.4M | L6 d408 h6 | 16B† | 70.53 | 40.91 | 1.1752 | | `run_32m_16b/ckpt.pt` (Path-B v1) | 31.4M | L6 d576 h9 | 16B‡ | **73.77** | **44.44** | 1.1128 | † crown = expanded 8.29B-token corpus (~1.9 epochs); board-recon **#14/74** (up from #18 via ARC). ‡ Path-B v1 = 32M at 16B tokens (ARC-MIX corpus): **breaks the 16M BLiMP ceiling** (70.5 -> 73.77) and lifts ARC to 44.44 -> board-recon **#11/74**. The 70.5 cap was 16M-specific, not absolute; more capacity + tokens moves both axes. ## Usage These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel` weights. The model class and a ready board-scoring harness are included in this repo: - `train_gpt_ref.py` — GPT definition (rebuild the GPT of the tabled shape, `load_state_dict`, trim logits to vocab 12288). - `glint_parity_eval.py` — the exact Glint board-scoring forward (256-token clip, raw log-prob) for BLiMP / ARC-Easy / WikiText-2. ```python import torch from tokenizers import Tokenizer tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288 ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu") state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py ``` **Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above): ```bash # crown 16M @ expanded corpus python train_gpt_ref.py --data-dir \ --n-layer 6 --n-embd 408 --n-head 6 --block 1024 \ --batch 64 --steps 244141 --lr 6e-4 --min-lr 6e-5 \ --vocab 12288 --dtype uint16 --seed 1337 # Path-B 32M: --n-embd 576 --n-head 9 (same vocab 12288 / uint16 / BPE tokenizer) ``` The `--vocab 12288 --dtype uint16` flags select BPE-12k over uint16 token bins. The script's byte-level defaults (`--vocab 256 --dtype uint8`) and its header comment reflect its origin as a **standard-GPT control** compared against an experimental **BDH (fast-weights)** architecture — the leaderboard models here are the standard causal transformer in BPE mode and do **not** use BDH. ## Training data [`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b) — ~5.40B BPE-12k tokens, English, decontaminated against the benchmark test sets. A broad high-quality mix: FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News. A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) feeds later runs. **Training budget and epochs.** 16M trained on 16B tokens *seen* is intentional, not a chart error. The leaderboard scores efficiency = quality at a fixed tiny size, so you over-train to squeeze max quality from frozen capacity. Chinchilla-optimal (about 20x params = 0.32B for 16M) minimizes compute-optimal loss, which is NOT the leaderboard objective; top models train many tokens-per-param too. 16B seen over 8.29B unique corpus = about 1.9 epochs (each token seen ~1.9x, under the 2x repeat-degradation limit; val-loss healthy, zero memorization). Note: the crown 16B point changed BOTH tokens and corpus (5.4B to 8.29B expanded), so on the token-scan chart it is marked separately (star + dashed) and is not a pure token step. BLiMP saturation holds regardless: crown BLiMP 70.53 is below even the token-only projection 71.6. ## Evaluation All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e. the board-comparable definitions: - **BLiMP** — 67 configs (train split), each sentence clipped to the first 256 tokens, raw sentence log-prob preference (`good > bad`), no length normalization. - **ARC-Easy** — test split, zero-shot, raw accuracy over `LL(question + choice) - LL(question)`. - **WikiText-2** — byte-normalized bits-per-byte (the board's `wiki` field is byte-scale; token-perplexity is tokenizer-dependent and not directly comparable across models). A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones. **Positioning (honest):** leaderboard ranks are **reconstruction estimates** — we reverse-engineered the board scoring formula and validated it (it reproduces a published reference model's public rank exactly), then applied it to our Glint-protocol metrics. Under that reconstruction the 16M@10B checkpoint sits around #18/74, the 32M baseline around #20/74, and the crown (16M @ expanded 8.3B) around #14/74. The Path-B 32M@16B checkpoint reaches around **#11/74**, with a full-budget 32M run projected toward the top-5. These are credible estimates, **not** confirmed entries; an official submission is required to confirm. ## Key findings (single-factor study) - **Tokens drive BLiMP, not size — up to a ceiling.** BLiMP climbs with tokens (~+1.8pp per doubling, 3.2B->10B) then **saturates at 16M's ~70.5 ceiling** (crown 16B: 70.53, +0.17 over 10B — flat, below the token-only projection 71.6); 16M->32M at matched 10B tokens also left BLiMP flat. Size does not move it; tokens stop moving it near the cap. - **Capacity + knowledge drive ARC.** 16M->32M at matched tokens lifted ARC-Easy +3.07pp. - **The BLiMP ceiling is size-specific, not absolute.** 16M saturates ~70.5 on tokens; a 32M model at 16B tokens reaches BLiMP 73.77 (Path-B v1) and keeps rising - capacity, not data, is the binding constraint at the top. - **ARC gains are capacity-gated.** ARC-density upweighting was null at 16M (40.87 vs 40.91) but the 32M model reached ARC 44.44; the same data helps only when the model has capacity to exploit it. - **Efficiency is size-bonus-weighted**, so the smallest model reaching a given raw score ranks highest; ARC is the binding lever toward the top, targeted next via knowledge distillation. ## Roadmap - Crown run (**done**): 16M @ expanded 8.29B corpus, 16B tokens -> BLiMP 70.53 / ARC 40.91, board-recon **#14/74**. Finding: BLiMP data-lever exhausted at 16M (~70.5 cap); ARC is the sole lever toward #1. - ARC boost: knowledge distillation from a strong in-house teacher (decontaminated), plus science-dense data. - Larger raw-score variant under evaluation. ## Limitations Base (not instruction-tuned) research models at 16-32M parameters, English-only. Expect limited factual knowledge and coherence; not intended for production use. ## Provenance Full dialectical record, evaluation artifacts and eval-protocol details in labvault `21_09_GoLLeM-v5-Skalowanie-Glint/` (see `90-Ewaluacja/EvalHarnessParity.md`). Trained on RunPod RTX 5090.