Maggio33 commited on
Commit
1dd2a62
·
verified ·
1 Parent(s): 75dffb0

Usage: reference in-repo train_gpt_ref.py and glint_parity_eval.py

Browse files
Files changed (1) hide show
  1. README.md +5 -4
README.md CHANGED
@@ -42,16 +42,17 @@ varying only tokens and model width.
42
  ## Usage
43
 
44
  These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
45
- weights. Architecture is inferable from tensor shapes and the table above.
 
 
 
46
 
47
  ```python
48
  import torch
49
  from tokenizers import Tokenizer
50
  tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
51
  ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
52
- state = ckpt.get("model", ckpt) # nanoGPT-style GPT state dict
53
- # Rebuild a GPT of the tabled shape (e.g. L6 d408 h6, block 1024, tied embeddings),
54
- # load_state_dict(state), trim logits to vocab 12288, run token ids [B, T].
55
  ```
56
 
57
  ## Training data
 
42
  ## Usage
43
 
44
  These are raw nanoGPT-lineage checkpoints (plain `torch` state dicts), **not** `transformers` `AutoModel`
45
+ weights. The model class and a ready board-scoring harness are included in this repo:
46
+
47
+ - `train_gpt_ref.py` — GPT definition (rebuild the GPT of the tabled shape, `load_state_dict`, trim logits to vocab 12288).
48
+ - `glint_parity_eval.py` — the exact Glint board-scoring forward (256-token clip, raw log-prob) for BLiMP / ARC-Easy / WikiText-2.
49
 
50
  ```python
51
  import torch
52
  from tokenizers import Tokenizer
53
  tok = Tokenizer.from_file("tokenizer.json") # BPE-12k, vocab 12288
54
  ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
55
+ state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
 
 
56
  ```
57
 
58
  ## Training data