provenance: large/context.md
Browse files- large/context.md +58 -0
large/context.md
ADDED
|
@@ -0,0 +1,58 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# large v1
|
| 2 |
+
|
| 3 |
+
- **preset:** large
|
| 4 |
+
- **training corpus:** data/large (see corpus_index.md / corpus_stats.txt)
|
| 5 |
+
- **trained:** 2026-06-18 13:58 (checkpoint mtime)
|
| 6 |
+
- **training wall-clock:** 208.2 min
|
| 7 |
+
- **final val loss:** 1.1111 (iter 23000)
|
| 8 |
+
- **parameters:** 49.72M | vocab 198
|
| 9 |
+
- **note:** v1 trained on the large corpus as of 2026-06-18 BEFORE the xlarge expansion: 1,085,579,228 chars / 248 authors / 22 categories / vocab 198 / best val 1.1111. The corpus_index.md + corpus_stats.txt in this dir have been CORRECTED to that true training set (the live data/large has since grown toward the xlarge ~2B target).
|
| 10 |
+
|
| 11 |
+
## Training hyperparameters
|
| 12 |
+
|
| 13 |
+
| param | value |
|
| 14 |
+
|---|---|
|
| 15 |
+
| block_size | 384 |
|
| 16 |
+
| n_embd | 640 |
|
| 17 |
+
| n_head | 10 |
|
| 18 |
+
| n_layer | 10 |
|
| 19 |
+
| dropout | 0.0 |
|
| 20 |
+
| batch_size | 16 |
|
| 21 |
+
| max_iters | 24000 |
|
| 22 |
+
| eval_interval | 1000 |
|
| 23 |
+
| eval_iters | 80 |
|
| 24 |
+
| learning_rate | 0.00025 |
|
| 25 |
+
| vocab_size | 198 |
|
| 26 |
+
|
| 27 |
+
## Validation curve
|
| 28 |
+
|
| 29 |
+
| step | train | val |
|
| 30 |
+
|---:|---:|---:|
|
| 31 |
+
| 0 | 5.5531 | 5.5478 |
|
| 32 |
+
| 1000 | 2.0925 | 2.0524 |
|
| 33 |
+
| 2000 | 1.5805 | 1.5329 |
|
| 34 |
+
| 3000 | 1.4392 | 1.3825 |
|
| 35 |
+
| 4000 | 1.3836 | 1.3366 |
|
| 36 |
+
| 5000 | 1.3361 | 1.2880 |
|
| 37 |
+
| 6000 | 1.3129 | 1.2693 |
|
| 38 |
+
| 7000 | 1.2897 | 1.2436 |
|
| 39 |
+
| 8000 | 1.2580 | 1.2296 |
|
| 40 |
+
| 9000 | 1.2388 | 1.2164 |
|
| 41 |
+
| 10000 | 1.2397 | 1.1977 |
|
| 42 |
+
| 11000 | 1.2146 | 1.1828 |
|
| 43 |
+
| 12000 | 1.2159 | 1.1787 |
|
| 44 |
+
| 13000 | 1.2004 | 1.1669 |
|
| 45 |
+
| 14000 | 1.1967 | 1.1636 |
|
| 46 |
+
| 15000 | 1.1915 | 1.1576 |
|
| 47 |
+
| 16000 | 1.1814 | 1.1457 |
|
| 48 |
+
| 17000 | 1.1793 | 1.1388 |
|
| 49 |
+
| 18000 | 1.1634 | 1.1461 |
|
| 50 |
+
| 19000 | 1.1557 | 1.1360 |
|
| 51 |
+
| 20000 | 1.1552 | 1.1365 |
|
| 52 |
+
| 21000 | 1.1573 | 1.1241 |
|
| 53 |
+
| 22000 | 1.1556 | 1.1260 |
|
| 54 |
+
| 23000 | 1.1388 | 1.1111 |
|
| 55 |
+
| 23999 | 1.1367 | 1.1189 |
|
| 56 |
+
|
| 57 |
+
## Reproducing the corpus
|
| 58 |
+
`corpus_index.md` lists every author and work in the training set; re-run `python -m corpus add-author` / `add-topic` per that index and `make finalize` to rebuild an equivalent corpus.
|