File size: 5,112 Bytes
8793375 c16e50a 8793375 c16e50a 8793375 68eeed2 c16e50a 8793375 ac68d22 c16e50a ac68d22 8793375 ac68d22 8793375 ac68d22 8793375 c16e50a 8793375 ac68d22 8793375 ac68d22 1dd2a62 8793375 ac68d22 1dd2a62 ac68d22 8793375 c16e50a ac68d22 c16e50a ac68d22 8793375 ac68d22 e1befe5 ac68d22 e1befe5 ac68d22 e1befe5 ac68d22 e1befe5 ac68d22 8793375 ac68d22 8793375 ac68d22 8793375 ac68d22 8793375 ac68d22 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 | ---
license: apache-2.0
language:
- en
library_name: pytorch
pipeline_tag: text-generation
tags:
- tiny-lm
- gpt
- nanogpt
- glint-tiny-ml-leaderboard
- english
datasets:
- SlayerLab/minimal-en-corpus-5b
---
# GoLLeM-v5 — Tiny English Language Models (16M-32M)
Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
(nanoGPT lineage) trained for the
[Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
This repository is a controlled single-factor scaling study: identical architecture/hyper-parameters/seed,
varying only tokens and model width.
## Model details
- **Architecture:** decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings.
- **Sizes:** 16M variant = 6 layers / d_model 408 / 6 heads (17.4M params); 32M variant = 6 layers / d_model 576 / 9 heads (31.4M params).
- **Context length:** 1024 tokens.
- **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
- **Training:** AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024 (65,536 tok/step), seed 1337, bf16 (RTX 5090).
## Checkpoints
| checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
|---|---|---|---|---|---|---|
| `bpe16m_3.2B/ckpt.pt` | 17.4M | L6 d408 h6 | 3.2B | 67.40 | 38.22 | 1.2161 |
| `bpe16m_6B/ckpt.pt` | 17.4M | L6 d408 h6 | 6B | 68.92 | 39.10 | 1.1943 |
| `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 |
| `bpe32m_baseline/ckpt.pt` | 31.4M | L6 d576 h9 | 10B | 70.08 | 42.59 | 1.124 |
## Usage
These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
weights. The model class and a ready board-scoring harness are included in this repo:
- `train_gpt_ref.py` — GPT definition (rebuild the GPT of the tabled shape, `load_state_dict`, trim logits to vocab 12288).
- `glint_parity_eval.py` — the exact Glint board-scoring forward (256-token clip, raw log-prob) for BLiMP / ARC-Easy / WikiText-2.
```python
import torch
from tokenizers import Tokenizer
tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
```
## Training data
[`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
— ~5.40B BPE-12k tokens, English, decontaminated against the benchmark test sets. A broad high-quality mix:
FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) feeds later runs.
## Evaluation
All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e. the board-comparable definitions:
- **BLiMP** — 67 configs (train split), each sentence clipped to the first 256 tokens, raw sentence log-prob preference (`good > bad`), no length normalization.
- **ARC-Easy** — test split, zero-shot, raw accuracy over `LL(question + choice) - LL(question)`.
- **WikiText-2** — byte-normalized bits-per-byte (the board's `wiki` field is byte-scale; token-perplexity is tokenizer-dependent and not directly comparable across models).
A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
**Positioning (honest):** leaderboard ranks are **reconstruction estimates** — we reverse-engineered the board
scoring formula and validated it (it reproduces a published reference model's public rank exactly), then applied it
to our Glint-protocol metrics. Under that reconstruction the 16M@10B checkpoint sits around #18/74 and the 32M
baseline around #20/74. These are credible estimates, **not** confirmed entries; an official submission is required to confirm.
## Key findings (single-factor study)
- **Tokens drive BLiMP, not size.** BLiMP keeps climbing with tokens (~+1.8pp per doubling, 3.2B->10B) without plateauing on the board protocol; 16M->32M at matched 10B tokens left BLiMP flat.
- **Capacity + knowledge drive ARC.** 16M->32M at matched tokens lifted ARC-Easy +3.07pp.
- **Efficiency is size-bonus-weighted**, so the smallest model reaching a given raw score ranks highest; ARC is the binding lever toward the top, targeted next via knowledge distillation.
## Roadmap
- Crown run: 16M @ expanded ~8.3B-token corpus (BLiMP lift from more/better data).
- ARC boost: knowledge distillation from a strong in-house teacher (decontaminated), plus science-dense data.
- Larger raw-score variant under evaluation.
## Limitations
Base (not instruction-tuned) research models at 16-32M parameters, English-only. Expect limited factual knowledge and
coherence; not intended for production use.
## Provenance
Full dialectical record, evaluation artifacts and eval-protocol details in labvault
`21_09_GoLLeM-v5-Skalowanie-Glint/` (see `90-Ewaluacja/EvalHarnessParity.md`). Trained on RunPod RTX 5090.
|