Maggio33 commited on
Commit
c16e50a
Β·
verified Β·
1 Parent(s): 8793375

Proper English model card: Glint-protocol metrics, honest positioning, corrected corpus/positions

Browse files
Files changed (1) hide show
  1. README.md +42 -83
README.md CHANGED
@@ -2,108 +2,67 @@
2
  license: apache-2.0
3
  language:
4
  - en
 
 
5
  tags:
6
  - tiny-lm
7
- - babylm
8
- - language-model
9
- - from-scratch
10
  - gpt
11
- library_name: pytorch
12
- pipeline_tag: text-generation
 
 
 
13
  ---
14
 
15
- # GoLLeM-v5 β€” English Tiny-LM Checkpoints
16
-
17
- Training checkpoints for **GoLLeM-v5**, a family of small English language models
18
- (β‰ˆ17M–31M parameters) trained **from scratch** on a curated, decontaminated English
19
- corpus. Built for the **sub-100M efficiency regime** and evaluated on the tiny-ML
20
- leaderboard tasks: **BLiMP**, **ARC-Easy**, and **WikiText-2**.
21
-
22
- Trained on a single RunPod **RTX 5090**. Redundant backups: this HF repo (durable) +
23
- the training pod (origin) + local disk.
24
-
25
- ## Architecture
26
-
27
- All 16M checkpoints share one decoder-only GPT (`GPT-ref`):
28
-
29
- | | |
30
- |---|---|
31
- | Layers / width / heads | **6 / 408 / 6** (β‰ˆ17.4M params) |
32
- | Context length | 1024 |
33
- | Vocabulary | 12288 (BPE-12k) |
34
- | Optimizer | AdamW, lr 6e-4 β†’ 6e-5 (cosine) |
35
- | Precision / seed | bf16 / 1337 |
36
-
37
- The 32M baseline is the same recipe at **L6 / d576 / h9 (β‰ˆ31.4M params)**.
38
-
39
- ## Checkpoints
40
-
41
- ### Token-scaling scan (16M, single-factor: identical arch/hypers, only the token budget changes)
42
-
43
- | file | tokens | epochs | BLiMP | ARC-Easy | WikiText-2 (BPB) |
44
- |------|-------:|-------:|------:|---------:|-----------------:|
45
- | `bpe16m_3.2B/ckpt.pt` | 3.2B | 1.2 | 71.54 | 39.39 | 1.2161 |
46
- | `bpe16m_6B/ckpt.pt` | 6B | 2.2 | 73.29 | 41.04 | 1.1943 |
47
- | `bpe16m_10B/ckpt.pt` | 10B | 3.7 | 73.43 | 41.67 | 1.1815 |
48
-
49
- Metrics above are from `lm-evaluation-harness` (BLiMP full, ARC-Easy, WikiText-2 bits-per-byte).
50
-
51
- **Finding:** BLiMP saturates around ~6B tokens for this 16M model (Ξ” 6β†’10B = +0.14);
52
- ARC-Easy keeps improving at 10B; BPB decreases monotonically.
53
-
54
- ### 32M size-isolation baseline
55
 
56
- | file | params | tokens |
57
- |------|-------:|-------:|
58
- | `bpe32m_baseline/ckpt.pt` | 31.4M (L6/d576/h9) | 10B |
 
59
 
60
- ### Leaderboard protocol (Glint) results
61
 
62
- The public leaderboard uses a stricter evaluation protocol (256-token clip, no BOS,
63
- train-split BLiMP, bare-prompt ARC-Easy raw accuracy, byte-normalized WikiText). Under
64
- that protocol our numbers are lower than the `lm-eval` numbers above (offset β‰ˆ 3.1pp
65
- BLiMP / 2.2pp ARC) β€” **always cite the Glint numbers for the board**:
66
 
67
- | model | tokens | BLiMP | ARC-Easy | efficiency rank |
68
- |-------|-------:|------:|---------:|:---------------:|
69
- | 16M | 10B | 70.36 | 39.52 | #18 / 74 |
70
- | 32M | 10B | 70.08 | 42.59 | #20 / 74 |
 
 
71
 
72
- The 32M model ranks *below* the 16M model on efficiency: the leaderboard's size bonus
73
- favors smaller models, and the extra ARC gain does not offset the reduced bonus β€” hence
74
- the crown effort stays at the 16M size.
 
 
75
 
76
- ## Tokenizer
77
 
78
- `tokenizer.json` β€” BPE, vocab 12288 ("BPE-12k"), shared by every v5 model. Locked after a
79
- 3/3 verdict against a byte-level baseline (BLiMP +2Οƒ ∧ WikiText-2 BPB Ξ”0.238 ∧ improved
80
- leaderboard efficiency).
81
 
82
  ## Training data
83
 
84
- - Base corpus: **`SlayerLab/minimal-en-corpus-5b`** β€” 5.40B unique BPE-12k tokens, a
85
- 15-source English mixture (FineWeb-Edu, DCLM, code, math, books, QA, science),
86
- decontaminated against the benchmark test sets.
87
- - An **expanded 8.29B-token** corpus (base + additional FineWeb-Edu-100BT + OpenStax
88
- science, all deduplicated and decontaminated) backs the in-progress 16M crown run.
89
 
90
- ## Usage
91
 
92
- ```python
93
- import torch
94
- ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
95
- sd = ckpt["model"] if "model" in ckpt else ckpt
96
- # GPT-ref: L6 d408 h6, vocab 12288, block 1024.
97
- # Model class + forward: train_gpt_ref.py. Trim logits to vocab 12288 (no padded vocab).
98
- ```
99
 
100
- ## Status / roadmap
101
 
102
- - βœ… 16M token-scaling scan (3.2 / 6 / 10B) and 32M size baseline.
103
- - πŸ”„ 16M @ expanded 8.29B corpus (16B-token "crown" run) β€” in progress.
104
- - ⏭ ARC-focused knowledge distillation (KARD) to lift ARC-Easy.
105
 
106
  ## Provenance
107
 
108
- Full experiment log and artifacts in the project vault
109
- (`21_09_GoLLeM-v5-Skalowanie-Glint/`). Backup date: 2026-09-22.
 
2
  license: apache-2.0
3
  language:
4
  - en
5
+ library_name: pytorch
6
+ pipeline_tag: text-generation
7
  tags:
8
  - tiny-lm
 
 
 
9
  - gpt
10
+ - nanogpt
11
+ - glint-tiny-ml-leaderboard
12
+ - english
13
+ datasets:
14
+ - SlayerLab/minimal-en-corpus-5b
15
  ---
16
 
17
+ # GoLLeM-v5 β€” Tiny English Language Models (16M-32M)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
18
 
19
+ A family of **sub-100M-parameter English language models** trained for the
20
+ [Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
21
+ GPT-style decoder (nanoGPT lineage), BPE-12k tokenizer, 1024-token context.
22
+ This repository holds training checkpoints from a controlled single-factor token-scaling study.
23
 
24
+ ## Models
25
 
26
+ All share: BPE-12k tokenizer (`tokenizer.json`, vocab 12288), architecture per size, AdamW (lr 6e-4 -> 6e-5 cosine), seed 1337, bf16.
 
 
 
27
 
28
+ | checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
29
+ |---|---|---|---|---|---|---|
30
+ | `bpe16m_3.2B/ckpt.pt` | 17.4M | L6 d408 h6 | 3.2B | 67.40 | 38.22 | 1.2161 |
31
+ | `bpe16m_6B/ckpt.pt` | 17.4M | L6 d408 h6 | 6B | 68.92 | 39.10 | 1.1943 |
32
+ | `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 |
33
+ | `bpe32m_baseline/ckpt.pt` | 31.4M | L6 d576 h9 | 10B | 70.08 | 42.59 | 1.124 |
34
 
35
+ **All metrics are computed with the Glint benchmark protocol** (`Glint-1.3/benchmark.py`):
36
+ BLiMP = 67 configs (train split), first-256-token clip, raw sentence log-prob preference;
37
+ ARC-Easy = test split, zero-shot, raw accuracy (`LL(q+choice) - LL(q)`);
38
+ WikiText-2 = byte-normalized bits-per-byte (board's `wiki` field is byte-scale, not tokenizer-token-PPL).
39
+ Note: a generic `lm-eval-harness` run scores BLiMP/ARC ~2-3pp higher than the Glint protocol; the numbers above are the **board-comparable** ones.
40
 
41
+ ## Key findings (single-factor study)
42
 
43
+ - **Tokens drive BLiMP, not size.** On the Glint protocol BLiMP keeps climbing with tokens (+1.8pp per doubling, 3.2B->10B) and does **not** plateau; going 16M->32M at matched 10B tokens left BLiMP flat (70.36 -> 70.08).
44
+ - **Size + knowledge drive ARC.** 16M->32M at matched tokens lifted ARC-Easy +3.07pp (39.52 -> 42.59).
45
+ - **Efficiency is size-bonus-weighted**, so the smallest model that reaches a given raw score ranks highest; ARC is the binding lever toward the top of the board (targeted via knowledge distillation, in progress).
46
 
47
  ## Training data
48
 
49
+ [`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
50
+ β€” ~5.40B BPE-12k tokens, English, decontaminated. Broad high-quality mix:
51
+ FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
52
+ A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) is used for later runs.
 
53
 
54
+ ## Positioning (honest)
55
 
56
+ Leaderboard positions are **reconstruction estimates**: we reverse-engineered and validated the board scoring
57
+ formula (it reproduces a published reference model's rank exactly) and applied it to our Glint-protocol metrics.
58
+ They are credible estimates, **not** confirmed board entries; an official submission is required to confirm.
 
 
 
 
59
 
60
+ ## Intended use & limitations
61
 
62
+ Research artifacts for small-LM scaling studies and leaderboard work. English-only, base (not instruction-tuned)
63
+ models at 16-32M parameters: expect limited factual knowledge and coherence. Not for production use.
 
64
 
65
  ## Provenance
66
 
67
+ Full dialectical record, evaluation artifacts and methodology: labvault `21_09_GoLLeM-v5-Skalowanie-Glint/`
68
+ (including `90-Ewaluacja/EvalHarnessParity.md` for the eval-protocol details). Trained on RunPod RTX 5090.