Maggio33 commited on
Commit
02a1315
·
verified ·
1 Parent(s): ff77f7c

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +13 -0
README.md CHANGED
@@ -61,6 +61,19 @@ ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
61
  state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
62
  ```
63
 
 
 
 
 
 
 
 
 
 
 
 
 
 
64
  ## Training data
65
 
66
  [`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)
 
61
  state = ckpt.get("model", ckpt) # load into the GPT from train_gpt_ref.py
62
  ```
63
 
64
+ **Reproducing training.** `train_gpt_ref.py` is a general nanoGPT-style causal transformer whose vocabulary and bin dtype are **CLI-parameterized**. These leaderboard checkpoints were trained in **BPE-12k mode** (`--vocab 12288 --dtype uint16`), not the script's byte-level defaults. Exact command (shape from the table above):
65
+
66
+ ```bash
67
+ # crown 16M @ expanded corpus
68
+ python train_gpt_ref.py --data-dir <corpus> \
69
+ --n-layer 6 --n-embd 408 --n-head 6 --block 1024 \
70
+ --batch 64 --steps 244141 --lr 6e-4 --min-lr 6e-5 \
71
+ --vocab 12288 --dtype uint16 --seed 1337
72
+ # Path-B 32M: --n-embd 576 --n-head 9 (same vocab 12288 / uint16 / BPE tokenizer)
73
+ ```
74
+
75
+ The `--vocab 12288 --dtype uint16` flags select BPE-12k over uint16 token bins. The script's byte-level defaults (`--vocab 256 --dtype uint8`) and its header comment reflect its origin as a **standard-GPT control** compared against an experimental **BDH (fast-weights)** architecture — the leaderboard models here are the standard causal transformer in BPE mode and do **not** use BDH.
76
+
77
  ## Training data
78
 
79
  [`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b)