Polished professional card: model-details, usage, evaluation-protocol, key-findings, roadmap, limitations
Browse files
README.md
CHANGED
|
@@ -16,14 +16,21 @@ datasets:
|
|
| 16 |
|
| 17 |
# GoLLeM-v5 — Tiny English Language Models (16M-32M)
|
| 18 |
|
| 19 |
-
|
|
|
|
| 20 |
[Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
|
| 21 |
-
|
| 22 |
-
|
| 23 |
|
| 24 |
-
##
|
| 25 |
|
| 26 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 27 |
|
| 28 |
| checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
|
| 29 |
|---|---|---|---|---|---|---|
|
|
@@ -32,56 +39,61 @@ All share: BPE-12k tokenizer (`tokenizer.json`, vocab 12288), architecture per s
|
|
| 32 |
| `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 |
|
| 33 |
| `bpe32m_baseline/ckpt.pt` | 31.4M | L6 d576 h9 | 10B | 70.08 | 42.59 | 1.124 |
|
| 34 |
|
| 35 |
-
|
| 36 |
-
BLiMP = 67 configs (train split), first-256-token clip, raw sentence log-prob preference;
|
| 37 |
-
ARC-Easy = test split, zero-shot, raw accuracy (`LL(q+choice) - LL(q)`);
|
| 38 |
-
WikiText-2 = byte-normalized bits-per-byte (board's `wiki` field is byte-scale, not tokenizer-token-PPL).
|
| 39 |
-
Note: a generic `lm-eval-harness` run scores BLiMP/ARC ~2-3pp higher than the Glint protocol; the numbers above are the **board-comparable** ones.
|
| 40 |
|
| 41 |
-
|
|
|
|
| 42 |
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 46 |
|
| 47 |
## Training data
|
| 48 |
|
| 49 |
[`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
|
| 50 |
-
— ~5.40B BPE-12k tokens, English, decontaminated.
|
| 51 |
FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
|
| 52 |
-
A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science)
|
| 53 |
|
| 54 |
-
##
|
| 55 |
|
| 56 |
-
|
| 57 |
-
`AutoModel` weights. Architecture is inferable from tensor shapes (embedding `[vocab, d_model]`,
|
| 58 |
-
layer count from keys) and the table above; the tokenizer is `tokenizers`-format.
|
| 59 |
|
| 60 |
-
``
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
|
|
|
|
|
|
|
|
|
| 69 |
|
| 70 |
-
|
| 71 |
-
|
|
|
|
| 72 |
|
| 73 |
-
##
|
| 74 |
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
|
| 79 |
-
##
|
| 80 |
|
| 81 |
-
|
| 82 |
-
|
| 83 |
|
| 84 |
## Provenance
|
| 85 |
|
| 86 |
-
Full dialectical record, evaluation artifacts and
|
| 87 |
-
(
|
|
|
|
| 16 |
|
| 17 |
# GoLLeM-v5 — Tiny English Language Models (16M-32M)
|
| 18 |
|
| 19 |
+
Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
|
| 20 |
+
(nanoGPT lineage) trained for the
|
| 21 |
[Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
|
| 22 |
+
This repository is a controlled single-factor scaling study: identical architecture/hyper-parameters/seed,
|
| 23 |
+
varying only tokens and model width.
|
| 24 |
|
| 25 |
+
## Model details
|
| 26 |
|
| 27 |
+
- **Architecture:** decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings.
|
| 28 |
+
- **Sizes:** 16M variant = 6 layers / d_model 408 / 6 heads (17.4M params); 32M variant = 6 layers / d_model 576 / 9 heads (31.4M params).
|
| 29 |
+
- **Context length:** 1024 tokens.
|
| 30 |
+
- **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
|
| 31 |
+
- **Training:** AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024 (65,536 tok/step), seed 1337, bf16 (RTX 5090).
|
| 32 |
+
|
| 33 |
+
## Checkpoints
|
| 34 |
|
| 35 |
| checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
|
| 36 |
|---|---|---|---|---|---|---|
|
|
|
|
| 39 |
| `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 |
|
| 40 |
| `bpe32m_baseline/ckpt.pt` | 31.4M | L6 d576 h9 | 10B | 70.08 | 42.59 | 1.124 |
|
| 41 |
|
| 42 |
+
## Usage
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
|
| 44 |
+
These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
|
| 45 |
+
weights. Architecture is inferable from tensor shapes and the table above.
|
| 46 |
|
| 47 |
+
```python
|
| 48 |
+
import torch
|
| 49 |
+
from tokenizers import Tokenizer
|
| 50 |
+
tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
|
| 51 |
+
ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
|
| 52 |
+
state = ckpt.get("model", ckpt) # nanoGPT-style GPT state dict
|
| 53 |
+
# Rebuild a GPT of the tabled shape (e.g. L6 d408 h6, block 1024, tied embeddings),
|
| 54 |
+
# load_state_dict(state), trim logits to vocab 12288, run token ids [B, T].
|
| 55 |
+
```
|
| 56 |
|
| 57 |
## Training data
|
| 58 |
|
| 59 |
[`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
|
| 60 |
+
— ~5.40B BPE-12k tokens, English, decontaminated against the benchmark test sets. A broad high-quality mix:
|
| 61 |
FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
|
| 62 |
+
A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) feeds later runs.
|
| 63 |
|
| 64 |
+
## Evaluation
|
| 65 |
|
| 66 |
+
All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e. the board-comparable definitions:
|
|
|
|
|
|
|
| 67 |
|
| 68 |
+
- **BLiMP** — 67 configs (train split), each sentence clipped to the first 256 tokens, raw sentence log-prob preference (`good > bad`), no length normalization.
|
| 69 |
+
- **ARC-Easy** — test split, zero-shot, raw accuracy over `LL(question + choice) - LL(question)`.
|
| 70 |
+
- **WikiText-2** — byte-normalized bits-per-byte (the board's `wiki` field is byte-scale; token-perplexity is tokenizer-dependent and not directly comparable across models).
|
| 71 |
+
|
| 72 |
+
A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
|
| 73 |
+
|
| 74 |
+
**Positioning (honest):** leaderboard ranks are **reconstruction estimates** — we reverse-engineered the board
|
| 75 |
+
scoring formula and validated it (it reproduces a published reference model's public rank exactly), then applied it
|
| 76 |
+
to our Glint-protocol metrics. Under that reconstruction the 16M@10B checkpoint sits around #18/74 and the 32M
|
| 77 |
+
baseline around #20/74. These are credible estimates, **not** confirmed entries; an official submission is required to confirm.
|
| 78 |
+
|
| 79 |
+
## Key findings (single-factor study)
|
| 80 |
|
| 81 |
+
- **Tokens drive BLiMP, not size.** BLiMP keeps climbing with tokens (~+1.8pp per doubling, 3.2B->10B) without plateauing on the board protocol; 16M->32M at matched 10B tokens left BLiMP flat.
|
| 82 |
+
- **Capacity + knowledge drive ARC.** 16M->32M at matched tokens lifted ARC-Easy +3.07pp.
|
| 83 |
+
- **Efficiency is size-bonus-weighted**, so the smallest model reaching a given raw score ranks highest; ARC is the binding lever toward the top, targeted next via knowledge distillation.
|
| 84 |
|
| 85 |
+
## Roadmap
|
| 86 |
|
| 87 |
+
- Crown run: 16M @ expanded ~8.3B-token corpus (BLiMP lift from more/better data).
|
| 88 |
+
- ARC boost: knowledge distillation from a strong in-house teacher (decontaminated), plus science-dense data.
|
| 89 |
+
- Larger raw-score variant under evaluation.
|
| 90 |
|
| 91 |
+
## Limitations
|
| 92 |
|
| 93 |
+
Base (not instruction-tuned) research models at 16-32M parameters, English-only. Expect limited factual knowledge and
|
| 94 |
+
coherence; not intended for production use.
|
| 95 |
|
| 96 |
## Provenance
|
| 97 |
|
| 98 |
+
Full dialectical record, evaluation artifacts and eval-protocol details in labvault
|
| 99 |
+
`21_09_GoLLeM-v5-Skalowanie-Glint/` (see `90-Ewaluacja/EvalHarnessParity.md`). Trained on RunPod RTX 5090.
|