card: add 64M Qwen3 load-recipe (maintainer audit load-safety)
Browse files
README.md
CHANGED
|
@@ -73,6 +73,36 @@ ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
|
|
| 73 |
state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
|
| 74 |
```
|
| 75 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 76 |
**Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above):
|
| 77 |
|
| 78 |
```bash
|
|
|
|
| 73 |
state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
|
| 74 |
```
|
| 75 |
|
| 76 |
+
**64M flagship (Qwen3-arch) — load from the checkpoint's embedded config.** The 64M is a **Qwen3-style decoder** (RoPE + SwiGLU + RMSNorm + QK-Norm + value residuals), *not* the nanoGPT of the 16M/32M. Its architecture is self-described in `ckpt["config"]`; rebuild from that (do **not** load it as the plain 16M/32M GPT — a plain-GPT loader mis-loads to silent wrong logits):
|
| 77 |
+
|
| 78 |
+
```python
|
| 79 |
+
import torch
|
| 80 |
+
from types import SimpleNamespace
|
| 81 |
+
from train_gpt_ref import GPT
|
| 82 |
+
ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
|
| 83 |
+
c = ck["config"] # rope/swiglu/rmsnorm/qk_norm/value_residual; n_head 9, L14, d576, block 1024
|
| 84 |
+
cfg = SimpleNamespace(**c)
|
| 85 |
+
m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
|
| 86 |
+
m.load_state_dict(ck["model"], strict=False) # tied head
|
| 87 |
+
m.eval()
|
| 88 |
+
```
|
| 89 |
+
|
| 90 |
+
The bundled `glint_parity_eval.py` is arch-aware (auto-detects `ckpt["config"]`) — run it directly on `v1_muon/ckpt_400k.pt` to reproduce the board numbers (BLiMP 77.84 / ARC-Easy 47.94 bare-prompt / eff 77.51).
|
| 91 |
+
|
| 92 |
+
**64M flagship (`v1_muon/ckpt_400k.pt`) — Qwen3-style arch, config embedded in the checkpoint.** Load with the cfg from `ckpt["config"]` (RoPE / SwiGLU / QK-Norm / value-residual, `n_head=9` / L14 / d576 / block 1024 are self-describing), **not** the legacy 5-arg nanoGPT path. `glint_parity_eval.py`'s loader is cfg-aware for both lineages.
|
| 93 |
+
|
| 94 |
+
```python
|
| 95 |
+
import torch
|
| 96 |
+
from types import SimpleNamespace
|
| 97 |
+
from train_gpt_ref import GPT
|
| 98 |
+
ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
|
| 99 |
+
c = ck["config"] # self-describing: vocab 12288, n_layer 14, n_embd 576, n_head 9, block 1024
|
| 100 |
+
cfg = SimpleNamespace(**c) # Qwen3 flags: rope θ100k / swiglu / qk_norm / value_residual
|
| 101 |
+
m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
|
| 102 |
+
m.load_state_dict(ck["model"], strict=False) # tied head.weight
|
| 103 |
+
m.eval()
|
| 104 |
+
```
|
| 105 |
+
|
| 106 |
**Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above):
|
| 107 |
|
| 108 |
```bash
|