Maggio33 commited on
Commit
b443fdb
·
verified ·
1 Parent(s): 7a7066f

card: add 64M Qwen3 load-recipe (maintainer audit load-safety)

Browse files
Files changed (1) hide show
  1. README.md +30 -0
README.md CHANGED
@@ -73,6 +73,36 @@ ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
73
  state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
74
  ```
75
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
76
  **Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above):
77
 
78
  ```bash
 
73
  state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
74
  ```
75
 
76
+ **64M flagship (Qwen3-arch) — load from the checkpoint's embedded config.** The 64M is a **Qwen3-style decoder** (RoPE + SwiGLU + RMSNorm + QK-Norm + value residuals), *not* the nanoGPT of the 16M/32M. Its architecture is self-described in `ckpt["config"]`; rebuild from that (do **not** load it as the plain 16M/32M GPT — a plain-GPT loader mis-loads to silent wrong logits):
77
+
78
+ ```python
79
+ import torch
80
+ from types import SimpleNamespace
81
+ from train_gpt_ref import GPT
82
+ ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
83
+ c = ck["config"] # rope/swiglu/rmsnorm/qk_norm/value_residual; n_head 9, L14, d576, block 1024
84
+ cfg = SimpleNamespace(**c)
85
+ m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
86
+ m.load_state_dict(ck["model"], strict=False) # tied head
87
+ m.eval()
88
+ ```
89
+
90
+ The bundled `glint_parity_eval.py` is arch-aware (auto-detects `ckpt["config"]`) — run it directly on `v1_muon/ckpt_400k.pt` to reproduce the board numbers (BLiMP 77.84 / ARC-Easy 47.94 bare-prompt / eff 77.51).
91
+
92
+ **64M flagship (`v1_muon/ckpt_400k.pt`) — Qwen3-style arch, config embedded in the checkpoint.** Load with the cfg from `ckpt["config"]` (RoPE / SwiGLU / QK-Norm / value-residual, `n_head=9` / L14 / d576 / block 1024 are self-describing), **not** the legacy 5-arg nanoGPT path. `glint_parity_eval.py`'s loader is cfg-aware for both lineages.
93
+
94
+ ```python
95
+ import torch
96
+ from types import SimpleNamespace
97
+ from train_gpt_ref import GPT
98
+ ck = torch.load("v1_muon/ckpt_400k.pt", map_location="cpu", weights_only=False)
99
+ c = ck["config"] # self-describing: vocab 12288, n_layer 14, n_embd 576, n_head 9, block 1024
100
+ cfg = SimpleNamespace(**c) # Qwen3 flags: rope θ100k / swiglu / qk_norm / value_residual
101
+ m = GPT(c["vocab"], c["n_layer"], c["n_embd"], c["n_head"], c["block"], cfg)
102
+ m.load_state_dict(ck["model"], strict=False) # tied head.weight
103
+ m.eval()
104
+ ```
105
+
106
  **Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above):
107
 
108
  ```bash