card: 64M #6 flagship, board-protocol verified (glint bare-prompt: eff 77.51 / ARC 47.94)
Browse files
README.md
CHANGED
|
@@ -14,7 +14,7 @@ datasets:
|
|
| 14 |
- SlayerLab/minimal-en-corpus-5b
|
| 15 |
---
|
| 16 |
|
| 17 |
-
# GoLLeM-v5 β Tiny English Language Models (16M-
|
| 18 |
|
| 19 |
Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
|
| 20 |
(nanoGPT lineage) trained for the
|
|
@@ -24,8 +24,8 @@ varying only tokens and model width.
|
|
| 24 |
|
| 25 |
## Model details
|
| 26 |
|
| 27 |
-
- **Architecture:** decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings.
|
| 28 |
-
- **Sizes:** 16M
|
| 29 |
- **Context length:** 1024 tokens.
|
| 30 |
- **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
|
| 31 |
- **Training:** AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024 (65,536 tok/step), seed 1337, bf16 (RTX 5090).
|
|
@@ -43,6 +43,7 @@ varying only tokens and model width.
|
|
| 43 |
| `run_149m/ckpt.pt` (scaling ref) | 149M | β | 10BΒ§ | 76.99 | 49.66 | 1.2052 |
|
| 44 |
| `run_32m_18b/ckpt.pt` (v1b slope-check) | 31.6M | L6 d576 h9 | 18BΒΆ | 72.38 | 44.70 | 1.3431 |
|
| 45 |
| `run_32m_muon/ckpt.pt` (Muon optimizer) | 31.6M | L6 d576 h9 | 16Bβ | 72.29 | 42.89 | 1.3866 |
|
|
|
|
| 46 |
|
| 47 |
β crown = expanded 8.29B-token corpus (~1.9 epochs). This is the published **16M board entry: confirmed #21** (eff 74.49 β 70.53 / 40.91 / byte_ppl 2.6746). Earlier recon estimated #20 (an optimistic #14 used a wrong wiki estimate before the exact byte_ppl correction); the official merge landed #21.
|
| 48 |
|
|
@@ -54,6 +55,8 @@ varying only tokens and model width.
|
|
| 54 |
|
| 55 |
β Muon optimizer = 32M at 16B tokens, Muon optimizer (muon-lr 0.02), on the expanded 8.29B corpus. BLiMP 72.29 / ARC 42.89 / byte_ppl 2.615. **Optimizer verdict = inconclusive (confounded design):** this run also changed corpus (expanded vs Path-B's arcmix), so the β1.48 BLiMP mixes optimizer *and* data and cannot isolate Muon. A clean Muon single-factor is deferred to the 64M A/B (Muon vs AdamW, corpus held fixed). Diagnostic run.
|
| 56 |
|
|
|
|
|
|
|
| 57 |
## Usage
|
| 58 |
|
| 59 |
These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
|
|
@@ -104,6 +107,8 @@ All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e
|
|
| 104 |
|
| 105 |
A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
|
| 106 |
|
|
|
|
|
|
|
| 107 |
**Positioning β CONFIRMED, on the board.** PR #76 was merged into the Glint Tiny-ML Leaderboard (2026-09-23), maintainer-verified (checkpoints loaded directly; params confirmed: 32M = 31,601,664, 16M = 17,449,344 deduped tied-embeddings; architecture matches `train_gpt_ref.py`, standard nanoGPT BPE-12288). Official standings: **#16 GoLLeM-v5 32M** (eff 75.51 β BLiMP 73.77 / ARC-Easy 44.44 / WikiText-2 byte_ppl 2.5386, 16B tok) and **#21 GoLLeM-v5 16M** (eff 74.49 β BLiMP 70.53 / ARC 40.91 / byte_ppl 2.6746). The board efficiency formula was reverse-engineered and then confirmed **line-for-line against the Space source** (reproduces the displayed eff exactly, 3/3 checked models to 2 decimals): `eff = mean(BLiMP, ARC-Easy, normalized-WikiText-2) Γ size-bonus`, where the size-bonus runs 1.0Γ (largest on board) to 1.5Γ (smallest) on a log-parameter scale. Our 32M carries a 1.065Γ size-bonus vs the #1's 1.013Γ β a ~5% efficiency edge at equal raw metrics.
|
| 108 |
|
| 109 |
## Key findings (single-factor study)
|
|
|
|
| 14 |
- SlayerLab/minimal-en-corpus-5b
|
| 15 |
---
|
| 16 |
|
| 17 |
+
# GoLLeM-v5 β Tiny English Language Models (16M-64M)
|
| 18 |
|
| 19 |
Research checkpoints of **sub-100M-parameter English language models**, GPT-style decoders
|
| 20 |
(nanoGPT lineage) trained for the
|
|
|
|
| 24 |
|
| 25 |
## Model details
|
| 26 |
|
| 27 |
+
- **Architecture:** 16M/32M = decoder-only Transformer (nanoGPT lineage), learned positional embeddings, tied input/output embeddings. **64M flagship = Qwen3-style decoder** (RoPE ΞΈ=100k, SwiGLU, RMSNorm, QK-Norm, value residuals), trained with the **Muon** optimizer.
|
| 28 |
+
- **Sizes:** 16M = 6 layers / d_model 408 / 6 heads (17.4M); 32M = 6 layers / d_model 576 / 9 heads (31.4M); **64M flagship = 14 layers / d_model 576 / 9 heads (62.9M)**.
|
| 29 |
- **Context length:** 1024 tokens.
|
| 30 |
- **Tokenizer:** BPE, vocab 12288 (`tokenizer.json`), shared across all checkpoints.
|
| 31 |
- **Training:** AdamW, lr 6e-4 -> 6e-5 (cosine), batch 64 x 1024 (65,536 tok/step), seed 1337, bf16 (RTX 5090).
|
|
|
|
| 43 |
| `run_149m/ckpt.pt` (scaling ref) | 149M | β | 10BΒ§ | 76.99 | 49.66 | 1.2052 |
|
| 44 |
| `run_32m_18b/ckpt.pt` (v1b slope-check) | 31.6M | L6 d576 h9 | 18BΒΆ | 72.38 | 44.70 | 1.3431 |
|
| 45 |
| `run_32m_muon/ckpt.pt` (Muon optimizer) | 31.6M | L6 d576 h9 | 16Bβ | 72.29 | 42.89 | 1.3866 |
|
| 46 |
+
| `v1_muon/ckpt_400k.pt` (**#6 flagship** β Qwen3+Muon+VR) | 62.9M | L14 d576 h9 | 13.1Bβ
| **77.84** | **47.94** | 1.012 |
|
| 47 |
|
| 48 |
β crown = expanded 8.29B-token corpus (~1.9 epochs). This is the published **16M board entry: confirmed #21** (eff 74.49 β 70.53 / 40.91 / byte_ppl 2.6746). Earlier recon estimated #20 (an optimistic #14 used a wrong wiki estimate before the exact byte_ppl correction); the official merge landed #21.
|
| 49 |
|
|
|
|
| 55 |
|
| 56 |
β Muon optimizer = 32M at 16B tokens, Muon optimizer (muon-lr 0.02), on the expanded 8.29B corpus. BLiMP 72.29 / ARC 42.89 / byte_ppl 2.615. **Optimizer verdict = inconclusive (confounded design):** this run also changed corpus (expanded vs Path-B's arcmix), so the β1.48 BLiMP mixes optimizer *and* data and cannot isolate Muon. A clean Muon single-factor is deferred to the 64M A/B (Muon vs AdamW, corpus held fixed). Diagnostic run.
|
| 57 |
|
| 58 |
+
β
64M flagship (v1 Muon) = 62.9M, **Qwen3-style decoder** (RoPE ΞΈ=100k + SwiGLU + RMSNorm + QK-Norm + value residuals), **Muon** optimizer (muon-lr 0.02, cosine), ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). **Board entry: #6 GoLLeM-v5 64M** (eff 77.51 β BLiMP 77.84 / ARC-Easy 47.94 / WikiText-2 byte_ppl 2.016, BPB 1.012). Recompute-verified **board-protocol** (Glint bare-prompt ARC β `LL(choice|q)` argmax, full BLiMP-67k, wiki byte_ppl, no-BOS β matching the maintainer's `glint_parity` harness), ckpt sha256 `59f982c1β¦`. A clean single-factor scale-up of the 32M Path-B recipe (same BPE-12288 tokenizer / arcmix data lineage; only size + the Qwen3+Muon+VR arch differ) that lifts eff 75.51 β 77.51. Also settles the clean **Muon verdict**: at 64M with corpus held fixed (Muon vs AdamW A/B), Muon wins on byte_ppl + BLiMP.
|
| 59 |
+
|
| 60 |
## Usage
|
| 61 |
|
| 62 |
These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
|
|
|
|
| 107 |
|
| 108 |
A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
|
| 109 |
|
| 110 |
+
**#6 β GoLLeM-v5 64M flagship (eff 77.51), board-protocol-verified (Glint bare-prompt ARC, 2026-09-24).** The 64M Muon model (Qwen3 arch + value residuals, ARC-MIX 9.42B) lands **#6 on the Glint Tiny-ML Leaderboard** β BLiMP 77.84 / ARC-Easy 47.94 / WikiText-2 byte_ppl 2.016 β behind only four 90β143M models and Glint-1.3 (982K, #5; razor-thin, eff 77.58 vs 77.51). It is the **strongest dense 64M entry** on the board, a **#16 β #6 jump** from the 32M. A clean single-factor scale-up (32M β 64M, same data/tokenizer lineage) plus the Qwen3+Muon+value-residual stack lifted eff 75.51 β 77.51. Numbers are recompute-verified **board-native** (bare-prompt ARC, matching the maintainer's `glint_parity` harness β **not** lm-eval), reproducible from the checkpoint.
|
| 111 |
+
|
| 112 |
**Positioning β CONFIRMED, on the board.** PR #76 was merged into the Glint Tiny-ML Leaderboard (2026-09-23), maintainer-verified (checkpoints loaded directly; params confirmed: 32M = 31,601,664, 16M = 17,449,344 deduped tied-embeddings; architecture matches `train_gpt_ref.py`, standard nanoGPT BPE-12288). Official standings: **#16 GoLLeM-v5 32M** (eff 75.51 β BLiMP 73.77 / ARC-Easy 44.44 / WikiText-2 byte_ppl 2.5386, 16B tok) and **#21 GoLLeM-v5 16M** (eff 74.49 β BLiMP 70.53 / ARC 40.91 / byte_ppl 2.6746). The board efficiency formula was reverse-engineered and then confirmed **line-for-line against the Space source** (reproduces the displayed eff exactly, 3/3 checked models to 2 decimals): `eff = mean(BLiMP, ARC-Easy, normalized-WikiText-2) Γ size-bonus`, where the size-bonus runs 1.0Γ (largest on board) to 1.5Γ (smallest) on a log-parameter scale. Our 32M carries a 1.065Γ size-bonus vs the #1's 1.013Γ β a ~5% efficiency edge at equal raw metrics.
|
| 113 |
|
| 114 |
## Key findings (single-factor study)
|