card: honest wiki correction (byte_ppl 2.372 / BPB 1.246 / eff ~75.9); maintainer re-benchmark PR#78; ARC-Easy 47.94 exact-match
Browse files
README.md
CHANGED
|
@@ -43,7 +43,7 @@ varying only tokens and model width.
|
|
| 43 |
| `run_149m/ckpt.pt` (scaling ref) | 149M | β | 10BΒ§ | 76.99 | 49.66 | 1.2052 |
|
| 44 |
| `run_32m_18b/ckpt.pt` (v1b slope-check) | 31.6M | L6 d576 h9 | 18BΒΆ | 72.38 | 44.70 | 1.3431 |
|
| 45 |
| `run_32m_muon/ckpt.pt` (Muon optimizer) | 31.6M | L6 d576 h9 | 16Bβ | 72.29 | 42.89 | 1.3866 |
|
| 46 |
-
| `v1_muon/ckpt_400k.pt` (**
|
| 47 |
|
| 48 |
β crown = expanded 8.29B-token corpus (~1.9 epochs). This is the published **16M board entry: confirmed #21** (eff 74.49 β 70.53 / 40.91 / byte_ppl 2.6746). Earlier recon estimated #20 (an optimistic #14 used a wrong wiki estimate before the exact byte_ppl correction); the official merge landed #21.
|
| 49 |
|
|
@@ -55,7 +55,7 @@ varying only tokens and model width.
|
|
| 55 |
|
| 56 |
β Muon optimizer = 32M at 16B tokens, Muon optimizer (muon-lr 0.02), on the expanded 8.29B corpus. BLiMP 72.29 / ARC 42.89 / byte_ppl 2.615. **Optimizer verdict = inconclusive (confounded design):** this run also changed corpus (expanded vs Path-B's arcmix), so the β1.48 BLiMP mixes optimizer *and* data and cannot isolate Muon. A clean Muon single-factor is deferred to the 64M A/B (Muon vs AdamW, corpus held fixed). Diagnostic run.
|
| 57 |
|
| 58 |
-
β
64M flagship (v1 Muon) = 62.9M, **Qwen3-style decoder** (RoPE ΞΈ=100k + SwiGLU + RMSNorm + QK-Norm + value residuals), **Muon** optimizer (muon-lr 0.02, cosine), ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). **
|
| 59 |
|
| 60 |
## Usage
|
| 61 |
|
|
@@ -73,22 +73,20 @@ ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
|
|
| 73 |
state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
|
| 74 |
```
|
| 75 |
|
| 76 |
-
**64M flagship (Qwen3-arch
|
| 77 |
|
| 78 |
```python
|
| 79 |
import torch
|
| 80 |
from types import SimpleNamespace
|
| 81 |
from train_gpt_ref import GPT
|
| 82 |
ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
|
| 83 |
-
c = ck["config"]
|
| 84 |
-
cfg = SimpleNamespace(**c)
|
| 85 |
m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
|
| 86 |
-
m.load_state_dict(ck["model"], strict=False) # tied head
|
| 87 |
m.eval()
|
| 88 |
```
|
| 89 |
|
| 90 |
-
The bundled `glint_parity_eval.py` is arch-aware (auto-detects `ckpt["config"]`) β run it directly on `v1_muon/ckpt_400k.pt` to reproduce the board numbers (BLiMP 77.84 / ARC-Easy 47.94 bare-prompt / eff 77.51).
|
| 91 |
-
|
| 92 |
**Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above):
|
| 93 |
|
| 94 |
```bash
|
|
@@ -123,7 +121,7 @@ All metrics use the **Glint benchmark protocol** (`Glint-1.3/benchmark.py`), i.e
|
|
| 123 |
|
| 124 |
A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
|
| 125 |
|
| 126 |
-
**
|
| 127 |
|
| 128 |
**Positioning β CONFIRMED, on the board.** PR #76 was merged into the Glint Tiny-ML Leaderboard (2026-09-23), maintainer-verified (checkpoints loaded directly; params confirmed: 32M = 31,601,664, 16M = 17,449,344 deduped tied-embeddings; architecture matches `train_gpt_ref.py`, standard nanoGPT BPE-12288). Official standings: **#16 GoLLeM-v5 32M** (eff 75.51 β BLiMP 73.77 / ARC-Easy 44.44 / WikiText-2 byte_ppl 2.5386, 16B tok) and **#21 GoLLeM-v5 16M** (eff 74.49 β BLiMP 70.53 / ARC 40.91 / byte_ppl 2.6746). The board efficiency formula was reverse-engineered and then confirmed **line-for-line against the Space source** (reproduces the displayed eff exactly, 3/3 checked models to 2 decimals): `eff = mean(BLiMP, ARC-Easy, normalized-WikiText-2) Γ size-bonus`, where the size-bonus runs 1.0Γ (largest on board) to 1.5Γ (smallest) on a log-parameter scale. Our 32M carries a 1.065Γ size-bonus vs the #1's 1.013Γ β a ~5% efficiency edge at equal raw metrics.
|
| 129 |
|
|
|
|
| 43 |
| `run_149m/ckpt.pt` (scaling ref) | 149M | β | 10BΒ§ | 76.99 | 49.66 | 1.2052 |
|
| 44 |
| `run_32m_18b/ckpt.pt` (v1b slope-check) | 31.6M | L6 d576 h9 | 18BΒΆ | 72.38 | 44.70 | 1.3431 |
|
| 45 |
| `run_32m_muon/ckpt.pt` (Muon optimizer) | 31.6M | L6 d576 h9 | 16Bβ | 72.29 | 42.89 | 1.3866 |
|
| 46 |
+
| `v1_muon/ckpt_400k.pt` (**flagship** β Qwen3+Muon+VR) | 62.9M | L14 d576 h9 | 13.1Bβ
| 75.99 | **47.94** | 1.246 |
|
| 47 |
|
| 48 |
β crown = expanded 8.29B-token corpus (~1.9 epochs). This is the published **16M board entry: confirmed #21** (eff 74.49 β 70.53 / 40.91 / byte_ppl 2.6746). Earlier recon estimated #20 (an optimistic #14 used a wrong wiki estimate before the exact byte_ppl correction); the official merge landed #21.
|
| 49 |
|
|
|
|
| 55 |
|
| 56 |
β Muon optimizer = 32M at 16B tokens, Muon optimizer (muon-lr 0.02), on the expanded 8.29B corpus. BLiMP 72.29 / ARC 42.89 / byte_ppl 2.615. **Optimizer verdict = inconclusive (confounded design):** this run also changed corpus (expanded vs Path-B's arcmix), so the β1.48 BLiMP mixes optimizer *and* data and cannot isolate Muon. A clean Muon single-factor is deferred to the 64M A/B (Muon vs AdamW, corpus held fixed). Diagnostic run.
|
| 57 |
|
| 58 |
+
β
64M flagship (v1 Muon) = 62.9M, **Qwen3-style decoder** (RoPE ΞΈ=100k + SwiGLU + RMSNorm + QK-Norm + value residuals), **Muon** optimizer (muon-lr 0.02, cosine), ARC-MIX 9.42B corpus, 400k steps = 13.1B tokens (~1.4 epochs). **Maintainer re-benchmark (PR #78, RTX 5090, strict load, sha256 `59f982c1β¦` matched): ARC-Easy 47.94 β an exact match to our number, the checkpoint-load + log-likelihood-scoring control β BLiMP 75.99, WikiText-2 byte_ppl 2.372 / BPB 1.246 β eff ~75.9.** An earlier version of this card cited eff 77.51 / byte_ppl 2.016 / BPB 1.012; that byte_ppl used a **wrong bytes-per-token conversion** (4.755 vs the canonical raw-text-bytes/tokens = 1,292,008 / 334,674 = 3.86 on WikiText-2-raw-v1 test), now corrected. Board-protocol (Glint bare-prompt ARC β `LL(choice|q)` argmax, full BLiMP-67k, wiki byte-normalized, no-BOS). A clean single-factor scale-up of the 32M Path-B recipe (same BPE-12288 tokenizer / arcmix data lineage; only size + the Qwen3+Muon+VR arch differ). Also settles the clean **Muon verdict**: at 64M with corpus held fixed (Muon vs AdamW A/B), Muon wins on byte_ppl + BLiMP.
|
| 59 |
|
| 60 |
## Usage
|
| 61 |
|
|
|
|
| 73 |
state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
|
| 74 |
```
|
| 75 |
|
| 76 |
+
**64M flagship (`v1_muon/ckpt_400k.pt`) β Qwen3-style arch, config embedded in the checkpoint.** Load with the cfg from `ckpt["config"]` (RoPE / SwiGLU / QK-Norm / value-residual, `n_head=9` / L14 / d576 / block 1024 are self-describing), **not** the legacy 5-arg nanoGPT path. `glint_parity_eval.py`'s loader is cfg-aware for both lineages.
|
| 77 |
|
| 78 |
```python
|
| 79 |
import torch
|
| 80 |
from types import SimpleNamespace
|
| 81 |
from train_gpt_ref import GPT
|
| 82 |
ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
|
| 83 |
+
c = ck["config"] # self-describing: vocab 12288, n_layer 14, n_embd 576, n_head 9, block 1024
|
| 84 |
+
cfg = SimpleNamespace(**c) # Qwen3 flags: rope ΞΈ100k / swiglu / qk_norm / value_residual
|
| 85 |
m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
|
| 86 |
+
m.load_state_dict(ck["model"], strict=False) # tied head.weight
|
| 87 |
m.eval()
|
| 88 |
```
|
| 89 |
|
|
|
|
|
|
|
| 90 |
**Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above):
|
| 91 |
|
| 92 |
```bash
|
|
|
|
| 121 |
|
| 122 |
A generic `lm-eval-harness` run scores BLiMP/ARC roughly 2-3pp higher than this protocol; the numbers here are the board-comparable ones.
|
| 123 |
|
| 124 |
+
**GoLLeM-v5 64M flagship (eff ~75.9), maintainer re-benchmarked (PR #78, 2026-09-24).** The 64M Muon model (Qwen3 arch + value residuals, ARC-MIX 9.42B) posts **ARC-Easy 47.94 β an exact match to our own number in the maintainer's independent RTX-5090 re-benchmark** (the checkpoint-load + log-likelihood-scoring control), with BLiMP 75.99 / WikiText-2 byte_ppl 2.372 β eff ~75.9. It is a strong **dense 64M entry**, a **#16 β 64M scale-up** from the 32M on the same data/tokenizer lineage plus the Qwen3+Muon+value-residual stack. An earlier version of this card cited eff 77.51 (byte_ppl 2.016) from a wrong bytes-per-token conversion β corrected here; better an honest ~75.9 than an inflated 77.5. Final board rank pending the maintainer's placement.
|
| 125 |
|
| 126 |
**Positioning β CONFIRMED, on the board.** PR #76 was merged into the Glint Tiny-ML Leaderboard (2026-09-23), maintainer-verified (checkpoints loaded directly; params confirmed: 32M = 31,601,664, 16M = 17,449,344 deduped tied-embeddings; architecture matches `train_gpt_ref.py`, standard nanoGPT BPE-12288). Official standings: **#16 GoLLeM-v5 32M** (eff 75.51 β BLiMP 73.77 / ARC-Easy 44.44 / WikiText-2 byte_ppl 2.5386, 16B tok) and **#21 GoLLeM-v5 16M** (eff 74.49 β BLiMP 70.53 / ARC 40.91 / byte_ppl 2.6746). The board efficiency formula was reverse-engineered and then confirmed **line-for-line against the Space source** (reproduces the displayed eff exactly, 3/3 checked models to 2 decimals): `eff = mean(BLiMP, ARC-Easy, normalized-WikiText-2) Γ size-bonus`, where the size-bonus runs 1.0Γ (largest on board) to 1.5Γ (smallest) on a log-parameter scale. Our 32M carries a 1.065Γ size-bonus vs the #1's 1.013Γ β a ~5% efficiency edge at equal raw metrics.
|
| 127 |
|