Maggio33 commited on
Commit
e1befe5
·
verified ·
1 Parent(s): c16e50a

Add Usage section (merge best-of-both cards)

Browse files
Files changed (1) hide show
  1. README.md +19 -0
README.md CHANGED
@@ -51,6 +51,25 @@ Note: a generic `lm-eval-harness` run scores BLiMP/ARC ~2-3pp higher than the Gl
51
  FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
52
  A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) is used for later runs.
53
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
54
  ## Positioning (honest)
55
 
56
  Leaderboard positions are **reconstruction estimates**: we reverse-engineered and validated the board scoring
 
51
  FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
52
  A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) is used for later runs.
53
 
54
+ ## Usage
55
+
56
+ These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers`
57
+ `AutoModel` weights. Architecture is inferable from tensor shapes (embedding `[vocab, d_model]`,
58
+ layer count from keys) and the table above; the tokenizer is `tokenizers`-format.
59
+
60
+ ```python
61
+ import torch
62
+ from tokenizers import Tokenizer
63
+ tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
64
+ ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
65
+ state = ckpt.get("model", ckpt) # nanoGPT-style GPT
66
+ # Rebuild a GPT with the shape from the table (e.g. L6 d408 h6, block 1024),
67
+ # load_state_dict(state), then run token ids [B, T] through the forward.
68
+ ```
69
+
70
+ For the exact board-scoring forward (Glint protocol: 256-token clip, raw log-prob), see
71
+ `glint_parity_eval.py` in the project staging.
72
+
73
  ## Positioning (honest)
74
 
75
  Leaderboard positions are **reconstruction estimates**: we reverse-engineered and validated the board scoring