Usage: reference in-repo train_gpt_ref.py and glint_parity_eval.py
Browse files
README.md
CHANGED
|
@@ -42,16 +42,17 @@ varying only tokens and model width.
|
|
| 42 |
## Usage
|
| 43 |
|
| 44 |
These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
|
| 45 |
-
weights.
|
|
|
|
|
|
|
|
|
|
| 46 |
|
| 47 |
```python
|
| 48 |
import torch
|
| 49 |
from tokenizers import Tokenizer
|
| 50 |
tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
|
| 51 |
ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
|
| 52 |
-
state = ckpt.get("model", ckpt) #
|
| 53 |
-
# Rebuild a GPT of the tabled shape (e.g. L6 d408 h6, block 1024, tied embeddings),
|
| 54 |
-
# load_state_dict(state), trim logits to vocab 12288, run token ids [B, T].
|
| 55 |
```
|
| 56 |
|
| 57 |
## Training data
|
|
|
|
| 42 |
## Usage
|
| 43 |
|
| 44 |
These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
|
| 45 |
+
weights. The model class and a ready board-scoring harness are included in this repo:
|
| 46 |
+
|
| 47 |
+
- `train_gpt_ref.py` — GPT definition (rebuild the GPT of the tabled shape, `load_state_dict`, trim logits to vocab 12288).
|
| 48 |
+
- `glint_parity_eval.py` — the exact Glint board-scoring forward (256-token clip, raw log-prob) for BLiMP / ARC-Easy / WikiText-2.
|
| 49 |
|
| 50 |
```python
|
| 51 |
import torch
|
| 52 |
from tokenizers import Tokenizer
|
| 53 |
tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
|
| 54 |
ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
|
| 55 |
+
state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
|
|
|
|
|
|
|
| 56 |
```
|
| 57 |
|
| 58 |
## Training data
|