File size: 10,116 Bytes
c14fa64 4736894 c14fa64 7ce15b7 c14fa64 7ce15b7 c14fa64 7ce15b7 c14fa64 7ce15b7 c14fa64 7ce15b7 c14fa64 7ce15b7 c14fa64 7ce15b7 c14fa64 7ce15b7 c14fa64 7ce15b7 c14fa64 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 | ---
license: mit
language:
- en
library_name: nanogpt
pipeline_tag: text-generation
tags:
- shakespeare
- nanogpt
- bpe
- gpt
- rope
- rmsnorm
- educational
datasets:
- shakespeare-complete-works
---
# Model Card β `shakespeare-nanogpt-3` (v3)
> A [sup computer](https://www.supcpu.com) release β a small language model studio. [Model page](https://www.supcpu.com/models/shakespeare-nanogpt-3/) Β· [monorepo](https://github.com/romellogoodman/sup-computer) (frozen code: [`projects/shakespeare/models/shakespeare-nanogpt-3/`](https://github.com/romellogoodman/sup-computer/tree/main/projects/shakespeare/models/shakespeare-nanogpt-3), tag `shakespeare-nanogpt-3`) Β· runs in your browser at [www.supcpu.com/model-player](https://www.supcpu.com/model-player/).
<div class="takeaways">
<p class="takeaways-label">Key takeaways</p>
<ul>
<li>The new series best: an enlarged early-modern-drama corpus + a <strong>1024-token, corpus-trained byte-level BPE</strong> + float32 training, reaching held-out <code>BPC 1.831</code> at just 11.02M params.</li>
<li>The headline is <strong>parameter efficiency</strong>: it beats the prior champion v2 (1.919 at 29.9M) and matches-or-beats a fresh float32 GPT-2-vocab control (1.843 at 29.9M) at ~1/3 the parameters.</li>
<li>The BPC edge over that fresh control is only β0.012 β <strong>within single-seed noise</strong>. The clean wins are the params efficiency and beating the prior champion; a multi-seed run is the stated next step.</li>
</ul>
</div>
The current best model in the [`shakespeare-nanogpt`](https://github.com/romellogoodman/sup-computer/blob/main/projects/shakespeare/README.md)
series: held-out BPC 1.831 at ~11.02M params, about a third the size of the
champion it beats. Where [v2](https://github.com/romellogoodman/sup-computer/blob/main/research-docs/model-cards/shakespeare-nanogpt-2.md) established the modern
architecture (RoPE, RMSNorm, bias-free) on the Complete Works with GPT-2 BPE,
**v3 changes the data and the tokenizer, not the architecture**: it trains a
small vocabulary *on the corpus itself*, enlarges that corpus with contemporary
early-modern drama, and trains in float32. This remains LLM-assisted research β
Claude as the *researcher* under human direction β not recursive
self-improvement.
> **Series note.** Successor to [`shakespeare-nanogpt-2`](https://github.com/romellogoodman/sup-computer/blob/main/research-docs/model-cards/shakespeare-nanogpt-2.md).
> All versions and the full story are in [`MODELS.md`](https://github.com/romellogoodman/sup-computer/blob/main/projects/shakespeare/MODELS.md);
> the scoreboard is [`leaderboard.md`](https://github.com/romellogoodman/sup-computer/blob/main/projects/shakespeare/leaderboard.md).
## Model details
| | |
|---|---|
| **Version / git tag** | `shakespeare-nanogpt-3` |
| **Origin** | LLM-assisted research, round 6 float32 re-run (`projects/shakespeare/runs/r6-fp32-bpe1k`) |
| **Architecture** | modern β RoPE, RMSNorm, bias-free (`core/nanogpt_core/model.py`, vendored into the frozen folder) |
| **Size** | ~11.02M params (6 layers Β· 6 heads Β· 384 embd Β· 256 context) |
| **Tokenizer** | 1024-vocab byte-level BPE, trained on the enlarged corpus (committed `tokenizer.json`; the `meta.pkl` seam, ADR-0012) |
| **Precision** | float32 (eliminates the MPS float16 large-vocab logit overflow that confounded round 5) |
| **Checkpoint** | `models/shakespeare-nanogpt-3/ckpt.pt` (weights not committed β rebuild below) |
| **Built on** | [nanoGPT](https://github.com/karpathy/nanoGPT) by Andrej Karpathy (MIT) |
| **Developed with** | Claude Fable 5 ([Claude Code](https://claude.com/claude-code)) as researcher, human oversight |
| **License** | MIT |
## Intended use
A **learning project** and a demonstration of LLM-assisted model development
(measured honestly, version over version). Given a few characters it continues
them in gibberish-but-convincingly-styled Early Modern English. v3's samples pick
up conventions from the wider corpus β speaker labels and the italic `_Name._`
stage-direction convention of the Marlowe/Webster editions.
**Out of scope:** real use of the text; any presentation of output as genuine
Shakespeare (or Marlowe, Jonson, etc.) or as fact. No instruction following, no
safety tuning. This is mimicry only.
## Training data
The **enlarged early-modern-drama corpus**: Shakespeare's Complete Works
(Gutenberg #100, ~5 MB) plus public-domain contemporary drama β Marlowe (*Doctor
Faustus*, both *Tamburlaine*s, *Edward II*, *The Jew of Malta*), Jonson
(*Volpone*, *The Alchemist*, *Every Man in His Humour*), Kyd (*The Spanish
Tragedy*), Webster (*The Duchess of Malfi*, *The White Devil*), and Dekker. Total
training text ~7.85M characters; the tokenizer is trained on the training split
only.
Crucially, the **held-out test set is unchanged**: the same fixed
250k-character Shakespeare slice (`projects/shakespeare/test.txt`) every version
in the series is scored on. It is excluded from training and never duplicated β
so v3's BPC stays directly comparable to v1 and v2 despite the larger, broader
training corpus. Enlarging the corpus also eliminated the overfit that defined
v2's rounds: validation loss fell monotonically instead of bottoming early.
## Training procedure
Trained with the vendored `train.py` on Apple Silicon (MPS), `dtype=float32`,
`lr=1e-3`, `block_size=256`, 2000 iterations, best-val checkpoint. Float32 is the
load-bearing change: round 5 hit a float16 instability on MPS (the CUDA-only
`GradScaler` is disabled there, so large-vocab logits overflowed and forced the
larger vocabularies to a crippled learning rate). Re-running every vocabulary at
float32 removed the overflow and let `bpe1k` vs the GPT-2 control finally be
compared apples-to-apples.
## Evaluation
Scored on the fixed held-out test in **bits-per-character (BPC)** β a
tokenizer-agnostic metric (total NLL of the test text Γ· its character count Γ·
ln 2), so char-level, GPT-2-BPE, and custom-BPE models are all directly
comparable. Lower is better.
| Model | Tokenizer | Params | Test BPC |
|-------|-----------|-------:|---------:|
| `shakespeare-nanogpt-1` (v1) | char (65) | 10.66M | 2.395 |
| `shakespeare-nanogpt-2` (v2) | GPT-2 BPE (50257) | 29.94M | 1.919 |
| r6 float32 GPT-2 control | GPT-2 BPE (50257) | 29.94M | 1.843 |
| **`shakespeare-nanogpt-3` (v3)** | **BPE (1024)** | **11.02M** | **1.831** |
| r6 float32 `bpe4k` (unreleased) | BPE (4096) | 12.19M | 1.813 |
Two clean results and one honest caveat:
- **Parameter efficiency (clean win).** v3 reaches equal-or-better BPC than the
50k-vocab models at ~1/3 the parameters. Of v2's 29.9M parameters, ~19.3M
*was* the GPT-2 embedding table; a 1024-vocab embedding is ~0.4M. That budget
buys capability instead of a giant vocabulary the Shakespeare domain never uses.
- **Beats the prior champion (clean win).** 1.831 < v2's 1.919 (β4.6%). All three
float32 models clear 1.919 β partly because float32 helps every vocabulary (the
GPT-2 control alone improves 1.919 β 1.843).
- **The edge over the fresh control is within noise (the caveat).** v3's β0.012
over the float32 GPT-2 control (1.831 vs 1.843) is smaller than the series' known
single-seed wobble; the `bpe4k` run is even lower (1.813) but likewise unreplicated.
Read the *vs-control* and *bpe1k-vs-bpe4k* gaps as **directionally promising, not
decisive.**
> **Chart to add (dataviz pipeline).** A grouped horizontal bar chart, *held-out
> BPC vs. parameter count* for the five rows above, would make the efficiency
> headline visible at a glance β v3 and the unreleased `bpe4k` sitting lowest on
> BPC while also furthest left on params, the two 29.9M GPT-2-vocab models to
> their right. Per repo convention this must be generated by
> [`dataviz/`](https://github.com/romellogoodman/sup-computer/blob/main/tools/dataviz/README.md) (add it to `build.py`), not
> hand-authored; it is described here rather than embedded because that build step
> is deferred.
## Limitations
- **Single-seed measurements.** Every score is one training run at a fixed seed
(`1337`), no variance estimate. v3's win over the *fresh float32 control* and its
gap to the unreleased `bpe4k` are within plausible run-to-run noise. The clear
results (params efficiency; beating the prior champion) do not depend on that
narrow margin. Multi-seed replication is the explicit next step β it is what
would turn "bpe1k is tied-best" into "bpe1k is best."
- **`bpe1k` was still improving.** In round 5 its validation loss had not
plateaued at 2000 iterations; more iterations would likely lower BPC further, so
1.831 is an under-trained floor, not a tuned optimum.
- **Domain-narrow test, broadened train.** The training corpus now includes
non-Shakespeare drama, but the test set is still pure Shakespeare β the metric
rewards Shakespeare mimicry specifically.
- **Still mimicry:** fluent early-modern *texture*, but no meaning, knowledge, or
factuality; no instruction following or safety tuning.
## How to use
The BPE steps need the Hugging Face `tokenizers` library, provided ad hoc:
```bash
# self-contained v3 folder (weights are gitignored β rebuild them)
cd projects/shakespeare/models/shakespeare-nanogpt-3
uv run --with tokenizers python prepare.py # downloads the enlarged corpus, encodes here
uv run python train.py # -> ./ckpt.pt (zero-arg run reproduces v3)
uv run --with tokenizers python eval.py # score on the shared held-out test (expect BPC ~1.831)
uv run --with tokenizers python sample.py --start="ROMEO:"
```
The 1024-token `tokenizer.json` is committed and never retrained β `prepare.py`
only re-encodes with it, pinning the exact vocabulary.
## Citation / credits
- nanoGPT by Andrej Karpathy (MIT) β model + training code.
- The Complete Works of Shakespeare and the contemporary early-modern drama
(Marlowe, Jonson, Kyd, Webster, Dekker) β all public domain (Project Gutenberg).
- LLM-assisted research run with Claude Fable 5 (Claude Code) as researcher, human oversight.
|