Maggio33 commited on
Commit
ac68d22
·
verified ·
1 Parent(s): e1befe5

Polished professional card: model-details, usage, evaluation-protocol, key-findings, roadmap, limitations

Browse files
Files changed (1) hide show
  1. README.md +52 -40
README.md CHANGED
@@ -16,14 +16,21 @@ datasets:
16
 
17
  # GoLLeM-v5 — Tiny English Language Models (16M-32M)
18
 
19
- A family of **sub-100M-parameter English language models** trained for the
 
20
  [Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
21
- GPT-style decoder (nanoGPT lineage), BPE-12k tokenizer, 1024-token context.
22
- This repository holds training checkpoints from a controlled single-factor token-scaling study.
23
 
24
- ## Models
25
 
26
- All share: BPE-12k tokenizer (`tokenizer.json`, vocab 12288), architecture per size, AdamW (lr 6e-4 -> 6e-5 cosine), seed 1337, bf16.
 
 
 
 
 
 
27
 
28
  | checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
29
  |---|---|---|---|---|---|---|
@@ -32,56 +39,61 @@ All share: BPE-12k tokenizer (`tokenizer.json`, vocab 12288), architecture per s
32
  | `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 |
33
  | `bpe32m_baseline/ckpt.pt` | 31.4M | L6 d576 h9 | 10B | 70.08 | 42.59 | 1.124 |
34
 
35
- **All metrics are computed with the Glint benchmark protocol** (`Glint-1.3/benchmark.py`):
36
- BLiMP = 67 configs (train split), first-256-token clip, raw sentence log-prob preference;
37
- ARC-Easy = test split, zero-shot, raw accuracy (`LL(q+choice) - LL(q)`);
38
- WikiText-2 = byte-normalized bits-per-byte (board's `wiki` field is byte-scale, not tokenizer-token-PPL).
39
- Note: a generic `lm-eval-harness` run scores BLiMP/ARC ~2-3pp higher than the Glint protocol; the numbers above are the **board-comparable** ones.
40
 
41
- ## Key findings (single-factor study)
 
42
 
43
- - **Tokens drive BLiMP, not size.** On the Glint protocol BLiMP keeps climbing with tokens (+1.8pp per doubling, 3.2B->10B) and does **not** plateau; going 16M->32M at matched 10B tokens left BLiMP flat (70.36 -> 70.08).
44
- - **Size + knowledge drive ARC.** 16M->32M at matched tokens lifted ARC-Easy +3.07pp (39.52 -> 42.59).
45
- - **Efficiency is size-bonus-weighted**, so the smallest model that reaches a given raw score ranks highest; ARC is the binding lever toward the top of the board (targeted via knowledge distillation, in progress).
 
 
 
 
 
 
46
 
47
  ## Training data
48
 
49
  [`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
50
- — ~5.40B BPE-12k tokens, English, decontaminated. Broad high-quality mix:
51
  FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
52
- A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) is used for later runs.
53
 
54
- ## Usage
55
 
56
- These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers`
57
- `AutoModel` weights. Architecture is inferable from tensor shapes (embedding `[vocab, d_model]`,
58
- layer count from keys) and the table above; the tokenizer is `tokenizers`-format.
59
 
60
- ```python
61
- import torch
62
- from tokenizers import Tokenizer
63
- tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
64
- ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
65
- state = ckpt.get("model", ckpt) # nanoGPT-style GPT
66
- # Rebuild a GPT with the shape from the table (e.g. L6 d408 h6, block 1024),
67
- # load_state_dict(state), then run token ids [B, T] through the forward.
68
- ```
 
 
 
69
 
70
- For the exact board-scoring forward (Glint protocol: 256-token clip, raw log-prob), see
71
- `glint_parity_eval.py` in the project staging.
 
72
 
73
- ## Positioning (honest)
74
 
75
- Leaderboard positions are **reconstruction estimates**: we reverse-engineered and validated the board scoring
76
- formula (it reproduces a published reference model's rank exactly) and applied it to our Glint-protocol metrics.
77
- They are credible estimates, **not** confirmed board entries; an official submission is required to confirm.
78
 
79
- ## Intended use & limitations
80
 
81
- Research artifacts for small-LM scaling studies and leaderboard work. English-only, base (not instruction-tuned)
82
- models at 16-32M parameters: expect limited factual knowledge and coherence. Not for production use.
83
 
84
  ## Provenance
85
 
86
- Full dialectical record, evaluation artifacts and methodology: labvault `21_09_GoLLeM-v5-Skalowanie-Glint/`
87
- (including `90-Ewaluacja/EvalHarnessParity.md` for the eval-protocol details). Trained on RunPod RTX 5090.
 
16
 
17
  # GoLLeM-v5 — Tiny English Language Models (16M-32M)
18
 
19
+ Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
20
+ (nanoGPT lineage) trained for the
21
  [Glint Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard).
22
+ This repository is a controlled single-factor scaling study: identical architecture/hyper-parameters/seed,
23
+ varying only tokens and model width.
24
 
25
+ ## Model details
26
 
27
+ - **Architecture:** decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings.
28
+ - **Sizes:** 16M variant = 6 layers / d_model 408 / 6 heads (17.4M params); 32M variant = 6 layers / d_model 576 / 9 heads (31.4M params).
29
+ - **Context length:** 1024 tokens.
30
+ - **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
31
+ - **Training:** AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024 (65,536 tok/step), seed 1337, bf16 (RTX 5090).
32
+
33
+ ## Checkpoints
34
 
35
  | checkpoint | params | shape | tokens | BLiMP | ARC-Easy | WikiText-2 BPB |
36
  |---|---|---|---|---|---|---|
 
39
  | `bpe16m_10B/ckpt.pt` | 17.4M | L6 d408 h6 | 10B | 70.36 | 39.52 | 1.1815 |
40
  | `bpe32m_baseline/ckpt.pt` | 31.4M | L6 d576 h9 | 10B | 70.08 | 42.59 | 1.124 |
41
 
42
+ ## Usage
 
 
 
 
43
 
44
+ These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
45
+ weights. Architecture is inferable from tensor shapes and the table above.
46
 
47
+ ```python
48
+ import torch
49
+ from tokenizers import Tokenizer
50
+ tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
51
+ ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
52
+ state = ckpt.get("model", ckpt) # nanoGPT-style GPT state dict
53
+ # Rebuild a GPT of the tabled shape (e.g. L6 d408 h6, block 1024, tied embeddings),
54
+ # load_state_dict(state), trim logits to vocab 12288, run token ids [B, T].
55
+ ```
56
 
57
  ## Training data
58
 
59
  [`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
60
+ — ~5.40B BPE-12k tokens, English, decontaminated against the benchmark test sets. A broad high-quality mix:
61
  FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
62
+ A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) feeds later runs.
63
 
64
+ ## Evaluation
65
 
66
+ All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e. the board-comparable definitions:
 
 
67
 
68
+ - **BLiMP** — 67 configs (train split), each sentence clipped to the first 256 tokens, raw sentence log-prob preference (`good > bad`), no length normalization.
69
+ - **ARC-Easy** — test split, zero-shot, raw accuracy over `LL(question + choice) - LL(question)`.
70
+ - **WikiText-2** — byte-normalized bits-per-byte (the board's `wiki` field is byte-scale; token-perplexity is tokenizer-dependent and not directly comparable across models).
71
+
72
+ A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
73
+
74
+ **Positioning (honest):** leaderboard ranks are **reconstruction estimates** — we reverse-engineered the board
75
+ scoring formula and validated it (it reproduces a published reference model's public rank exactly), then applied it
76
+ to our Glint-protocol metrics. Under that reconstruction the 16M@10B checkpoint sits around #18/74 and the 32M
77
+ baseline around #20/74. These are credible estimates, **not** confirmed entries; an official submission is required to confirm.
78
+
79
+ ## Key findings (single-factor study)
80
 
81
+ - **Tokens drive BLiMP, not size.** BLiMP keeps climbing with tokens (~+1.8pp per doubling, 3.2B->10B) without plateauing on the board protocol; 16M->32M at matched 10B tokens left BLiMP flat.
82
+ - **Capacity + knowledge drive ARC.** 16M->32M at matched tokens lifted ARC-Easy +3.07pp.
83
+ - **Efficiency is size-bonus-weighted**, so the smallest model reaching a given raw score ranks highest; ARC is the binding lever toward the top, targeted next via knowledge distillation.
84
 
85
+ ## Roadmap
86
 
87
+ - Crown run: 16M @ expanded ~8.3B-token corpus (BLiMP lift from more/better data).
88
+ - ARC boost: knowledge distillation from a strong in-house teacher (decontaminated), plus science-dense data.
89
+ - Larger raw-score variant under evaluation.
90
 
91
+ ## Limitations
92
 
93
+ Base (not instruction-tuned) research models at 16-32M parameters, English-only. Expect limited factual knowledge and
94
+ coherence; not intended for production use.
95
 
96
  ## Provenance
97
 
98
+ Full dialectical record, evaluation artifacts and eval-protocol details in labvault
99
+ `21_09_GoLLeM-v5-Skalowanie-Glint/` (see `90-Ewaluacja/EvalHarnessParity.md`). Trained on RunPod RTX 5090.