Add Usage section (merge best-of-both cards)
Browse files
README.md
CHANGED
|
@@ -51,6 +51,25 @@ Note: a generic `lm-eval-harness` run scores BLiMP/ARC ~2-3pp higher than the Gl
|
|
| 51 |
FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
|
| 52 |
A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) is used for later runs.
|
| 53 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 54 |
## Positioning (honest)
|
| 55 |
|
| 56 |
Leaderboard positions are **reconstruction estimates**: we reverse-engineered and validated the board scoring
|
|
|
|
| 51 |
FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News.
|
| 52 |
A decontaminated expansion to ~8.3B tokens (added FineWeb-Edu + OpenStax science) is used for later runs.
|
| 53 |
|
| 54 |
+
## Usage
|
| 55 |
+
|
| 56 |
+
These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers`
|
| 57 |
+
`AutoModel` weights. Architecture is inferable from tensor shapes (embedding `[vocab, d_model]`,
|
| 58 |
+
layer count from keys) and the table above; the tokenizer is `tokenizers`-format.
|
| 59 |
+
|
| 60 |
+
```python
|
| 61 |
+
import torch
|
| 62 |
+
from tokenizers import Tokenizer
|
| 63 |
+
tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
|
| 64 |
+
ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
|
| 65 |
+
state = ckpt.get("model", ckpt) # nanoGPT-style GPT
|
| 66 |
+
# Rebuild a GPT with the shape from the table (e.g. L6 d408 h6, block 1024),
|
| 67 |
+
# load_state_dict(state), then run token ids [B, T] through the forward.
|
| 68 |
+
```
|
| 69 |
+
|
| 70 |
+
For the exact board-scoring forward (Glint protocol: 256-token clip, raw log-prob), see
|
| 71 |
+
`glint_parity_eval.py` in the project staging.
|
| 72 |
+
|
| 73 |
## Positioning (honest)
|
| 74 |
|
| 75 |
Leaderboard positions are **reconstruction estimates**: we reverse-engineered and validated the board scoring
|