File size: 10,116 Bytes
c14fa64
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4736894
c14fa64
 
 
 
7ce15b7
 
 
c14fa64
 
 
 
7ce15b7
 
 
 
 
 
 
 
c14fa64
 
 
 
 
 
 
 
 
 
 
 
 
7ce15b7
 
c14fa64
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7ce15b7
c14fa64
 
 
 
 
 
 
 
7ce15b7
 
c14fa64
 
 
 
 
7ce15b7
c14fa64
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7ce15b7
c14fa64
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7ce15b7
c14fa64
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7ce15b7
c14fa64
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
---
license: mit
language:
  - en
library_name: nanogpt
pipeline_tag: text-generation
tags:
  - shakespeare
  - nanogpt
  - bpe
  - gpt
  - rope
  - rmsnorm
  - educational
datasets:
  - shakespeare-complete-works
---

# Model Card β€” `shakespeare-nanogpt-3` (v3)

> A [sup computer](https://www.supcpu.com) release β€” a small language model studio. [Model page](https://www.supcpu.com/models/shakespeare-nanogpt-3/) Β· [monorepo](https://github.com/romellogoodman/sup-computer) (frozen code: [`projects/shakespeare/models/shakespeare-nanogpt-3/`](https://github.com/romellogoodman/sup-computer/tree/main/projects/shakespeare/models/shakespeare-nanogpt-3), tag `shakespeare-nanogpt-3`) Β· runs in your browser at [www.supcpu.com/model-player](https://www.supcpu.com/model-player/).

<div class="takeaways">
<p class="takeaways-label">Key takeaways</p>
<ul>
<li>The new series best: an enlarged early-modern-drama corpus + a <strong>1024-token, corpus-trained byte-level BPE</strong> + float32 training, reaching held-out <code>BPC 1.831</code> at just 11.02M params.</li>
<li>The headline is <strong>parameter efficiency</strong>: it beats the prior champion v2 (1.919 at 29.9M) and matches-or-beats a fresh float32 GPT-2-vocab control (1.843 at 29.9M) at ~1/3 the parameters.</li>
<li>The BPC edge over that fresh control is only βˆ’0.012 β€” <strong>within single-seed noise</strong>. The clean wins are the params efficiency and beating the prior champion; a multi-seed run is the stated next step.</li>
</ul>
</div>

The current best model in the [`shakespeare-nanogpt`](https://github.com/romellogoodman/sup-computer/blob/main/projects/shakespeare/README.md)
series: held-out BPC 1.831 at ~11.02M params, about a third the size of the
champion it beats. Where [v2](https://github.com/romellogoodman/sup-computer/blob/main/research-docs/model-cards/shakespeare-nanogpt-2.md) established the modern
architecture (RoPE, RMSNorm, bias-free) on the Complete Works with GPT-2 BPE,
**v3 changes the data and the tokenizer, not the architecture**: it trains a
small vocabulary *on the corpus itself*, enlarges that corpus with contemporary
early-modern drama, and trains in float32. This remains LLM-assisted research β€”
Claude as the *researcher* under human direction β€” not recursive
self-improvement.

> **Series note.** Successor to [`shakespeare-nanogpt-2`](https://github.com/romellogoodman/sup-computer/blob/main/research-docs/model-cards/shakespeare-nanogpt-2.md).
> All versions and the full story are in [`MODELS.md`](https://github.com/romellogoodman/sup-computer/blob/main/projects/shakespeare/MODELS.md);
> the scoreboard is [`leaderboard.md`](https://github.com/romellogoodman/sup-computer/blob/main/projects/shakespeare/leaderboard.md).

## Model details

| | |
|---|---|
| **Version / git tag** | `shakespeare-nanogpt-3` |
| **Origin** | LLM-assisted research, round 6 float32 re-run (`projects/shakespeare/runs/r6-fp32-bpe1k`) |
| **Architecture** | modern β€” RoPE, RMSNorm, bias-free (`core/nanogpt_core/model.py`, vendored into the frozen folder) |
| **Size** | ~11.02M params (6 layers Β· 6 heads Β· 384 embd Β· 256 context) |
| **Tokenizer** | 1024-vocab byte-level BPE, trained on the enlarged corpus (committed `tokenizer.json`; the `meta.pkl` seam, ADR-0012) |
| **Precision** | float32 (eliminates the MPS float16 large-vocab logit overflow that confounded round 5) |
| **Checkpoint** | `models/shakespeare-nanogpt-3/ckpt.pt` (weights not committed β€” rebuild below) |
| **Built on** | [nanoGPT](https://github.com/karpathy/nanoGPT) by Andrej Karpathy (MIT) |
| **Developed with** | Claude Fable 5 ([Claude Code](https://claude.com/claude-code)) as researcher, human oversight |
| **License** | MIT |

## Intended use

A **learning project** and a demonstration of LLM-assisted model development
(measured honestly, version over version). Given a few characters it continues
them in gibberish-but-convincingly-styled Early Modern English. v3's samples pick
up conventions from the wider corpus β€” speaker labels and the italic `_Name._`
stage-direction convention of the Marlowe/Webster editions.

**Out of scope:** real use of the text; any presentation of output as genuine
Shakespeare (or Marlowe, Jonson, etc.) or as fact. No instruction following, no
safety tuning. This is mimicry only.

## Training data

The **enlarged early-modern-drama corpus**: Shakespeare's Complete Works
(Gutenberg #100, ~5 MB) plus public-domain contemporary drama β€” Marlowe (*Doctor
Faustus*, both *Tamburlaine*s, *Edward II*, *The Jew of Malta*), Jonson
(*Volpone*, *The Alchemist*, *Every Man in His Humour*), Kyd (*The Spanish
Tragedy*), Webster (*The Duchess of Malfi*, *The White Devil*), and Dekker. Total
training text ~7.85M characters; the tokenizer is trained on the training split
only.

Crucially, the **held-out test set is unchanged**: the same fixed
250k-character Shakespeare slice (`projects/shakespeare/test.txt`) every version
in the series is scored on. It is excluded from training and never duplicated β€”
so v3's BPC stays directly comparable to v1 and v2 despite the larger, broader
training corpus. Enlarging the corpus also eliminated the overfit that defined
v2's rounds: validation loss fell monotonically instead of bottoming early.

## Training procedure

Trained with the vendored `train.py` on Apple Silicon (MPS), `dtype=float32`,
`lr=1e-3`, `block_size=256`, 2000 iterations, best-val checkpoint. Float32 is the
load-bearing change: round 5 hit a float16 instability on MPS (the CUDA-only
`GradScaler` is disabled there, so large-vocab logits overflowed and forced the
larger vocabularies to a crippled learning rate). Re-running every vocabulary at
float32 removed the overflow and let `bpe1k` vs the GPT-2 control finally be
compared apples-to-apples.

## Evaluation

Scored on the fixed held-out test in **bits-per-character (BPC)** β€” a
tokenizer-agnostic metric (total NLL of the test text Γ· its character count Γ·
ln 2), so char-level, GPT-2-BPE, and custom-BPE models are all directly
comparable. Lower is better.

| Model | Tokenizer | Params | Test BPC |
|-------|-----------|-------:|---------:|
| `shakespeare-nanogpt-1` (v1) | char (65) | 10.66M | 2.395 |
| `shakespeare-nanogpt-2` (v2) | GPT-2 BPE (50257) | 29.94M | 1.919 |
| r6 float32 GPT-2 control | GPT-2 BPE (50257) | 29.94M | 1.843 |
| **`shakespeare-nanogpt-3` (v3)** | **BPE (1024)** | **11.02M** | **1.831** |
| r6 float32 `bpe4k` (unreleased) | BPE (4096) | 12.19M | 1.813 |

Two clean results and one honest caveat:

- **Parameter efficiency (clean win).** v3 reaches equal-or-better BPC than the
  50k-vocab models at ~1/3 the parameters. Of v2's 29.9M parameters, ~19.3M
  *was* the GPT-2 embedding table; a 1024-vocab embedding is ~0.4M. That budget
  buys capability instead of a giant vocabulary the Shakespeare domain never uses.
- **Beats the prior champion (clean win).** 1.831 < v2's 1.919 (βˆ’4.6%). All three
  float32 models clear 1.919 β€” partly because float32 helps every vocabulary (the
  GPT-2 control alone improves 1.919 β†’ 1.843).
- **The edge over the fresh control is within noise (the caveat).** v3's βˆ’0.012
  over the float32 GPT-2 control (1.831 vs 1.843) is smaller than the series' known
  single-seed wobble; the `bpe4k` run is even lower (1.813) but likewise unreplicated.
  Read the *vs-control* and *bpe1k-vs-bpe4k* gaps as **directionally promising, not
  decisive.**

> **Chart to add (dataviz pipeline).** A grouped horizontal bar chart, *held-out
> BPC vs. parameter count* for the five rows above, would make the efficiency
> headline visible at a glance β€” v3 and the unreleased `bpe4k` sitting lowest on
> BPC while also furthest left on params, the two 29.9M GPT-2-vocab models to
> their right. Per repo convention this must be generated by
> [`dataviz/`](https://github.com/romellogoodman/sup-computer/blob/main/tools/dataviz/README.md) (add it to `build.py`), not
> hand-authored; it is described here rather than embedded because that build step
> is deferred.

## Limitations

- **Single-seed measurements.** Every score is one training run at a fixed seed
  (`1337`), no variance estimate. v3's win over the *fresh float32 control* and its
  gap to the unreleased `bpe4k` are within plausible run-to-run noise. The clear
  results (params efficiency; beating the prior champion) do not depend on that
  narrow margin. Multi-seed replication is the explicit next step β€” it is what
  would turn "bpe1k is tied-best" into "bpe1k is best."
- **`bpe1k` was still improving.** In round 5 its validation loss had not
  plateaued at 2000 iterations; more iterations would likely lower BPC further, so
  1.831 is an under-trained floor, not a tuned optimum.
- **Domain-narrow test, broadened train.** The training corpus now includes
  non-Shakespeare drama, but the test set is still pure Shakespeare β€” the metric
  rewards Shakespeare mimicry specifically.
- **Still mimicry:** fluent early-modern *texture*, but no meaning, knowledge, or
  factuality; no instruction following or safety tuning.

## How to use

The BPE steps need the Hugging Face `tokenizers` library, provided ad hoc:

```bash
# self-contained v3 folder (weights are gitignored β€” rebuild them)
cd projects/shakespeare/models/shakespeare-nanogpt-3
uv run --with tokenizers python prepare.py   # downloads the enlarged corpus, encodes here
uv run python train.py                        # -> ./ckpt.pt (zero-arg run reproduces v3)
uv run --with tokenizers python eval.py       # score on the shared held-out test (expect BPC ~1.831)
uv run --with tokenizers python sample.py --start="ROMEO:"
```

The 1024-token `tokenizer.json` is committed and never retrained β€” `prepare.py`
only re-encodes with it, pinning the exact vocabulary.

## Citation / credits

- nanoGPT by Andrej Karpathy (MIT) β€” model + training code.
- The Complete Works of Shakespeare and the contemporary early-modern drama
  (Marlowe, Jonson, Kyd, Webster, Dekker) β€” all public domain (Project Gutenberg).
- LLM-assisted research run with Claude Fable 5 (Claude Code) as researcher, human oversight.