mellogood commited on
Commit
fd9aedf
Β·
verified Β·
1 Parent(s): 979e51a

release: daydream-chess-nanogpt-micro-1 (checkpoint + ONNX + tokenizer + model card)

Browse files
README.md ADDED
@@ -0,0 +1,133 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ language:
4
+ - en
5
+ library_name: nanogpt
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - daydream
9
+ - nanogpt
10
+ - char-level
11
+ - gpt
12
+ - chess
13
+ - uci
14
+ - minichess
15
+ ---
16
+
17
+ # Model Card β€” `daydream-chess-nanogpt-micro-1` (v1, Micro)
18
+
19
+ > A [sup computer](https://supcpu.romellogoodman.com) release β€” a small language model studio. [Model page](https://supcpu.romellogoodman.com/models/daydream-chess-nanogpt-micro-1/) Β· [monorepo](https://github.com/romellogoodman/sup-computer) (frozen code: [`projects/daydream/models/daydream-chess-nanogpt-micro-1/`](https://github.com/romellogoodman/sup-computer/tree/main/projects/daydream/models/daydream-chess-nanogpt-micro-1), tag `daydream-chess-nanogpt-micro-1`) Β· runs in your browser at [supcpu.romellogoodman.com/model-player](https://supcpu.romellogoodman.com/model-player/).
20
+
21
+
22
+ <div class="takeaways">
23
+ <p class="takeaways-label">Key takeaways</p>
24
+ <ul>
25
+ <li>A <strong>0.79M-param</strong> char-level GPT trained entirely on <strong>synthetic self-play</strong> β€” no human corpus exists for 5Γ—5 Gardner minichess, so all 4,135 training games came from two Fairy-Stockfish instances playing each other.</li>
26
+ <li>Fixed-depth engine self-play is <strong>fully deterministic</strong> on its own β€” the first generation attempt produced identical games every time. Fixed by randomizing opening plies (sourced from the engine's own legal-move list) before search takes over.</li>
27
+ <li><strong>100% clean completion, 39.2% legal-move rate</strong> on first try β€” slightly higher than the <a href="daydream-chess-nanogpt-1.md">Regular</a> tier's 35.3%, consistent with a smaller board being an easier legality problem to learn, though the corpora and vocab sizes differ too much to call it a controlled comparison.</li>
28
+ <li>Smallest tier in the three-board <a href="../../projects/daydream/README.md">daydream</a> family β€” 5Γ—5 is the smallest board that can hold one of every standard chess piece, which is why Micro uses Gardner's real, balance-tested arrangement rather than an invented one.</li>
29
+ </ul>
30
+ </div>
31
+
32
+ The smallest tier in the [`daydream`](https://github.com/romellogoodman/sup-computer/blob/main/projects/daydream/README.md)
33
+ family: a chess-move GPT trained on **Gardner minichess**, a real 5Γ—5 chess
34
+ variant β€” one each of King/Queen/Rook/Bishop/Knight per side, five pawns.
35
+ Same mechanic as the rest of the series: legal moves snap into focus,
36
+ illegal moves render as dim near-misses instead of being discarded.
37
+
38
+ > **A smaller board means a smaller book to memorize.** The animating
39
+ > thesis behind the daydream series is that repetition (opening theory,
40
+ > memorized lines) is where a model is most "in focus" and least
41
+ > interesting. Micro tests the far end of that: with only 25 squares and 6
42
+ > non-pawn pieces per side, there's very little room for memorized
43
+ > structure at all β€” almost everything the model does here, it has to
44
+ > generalize from a comparatively tiny, self-play-only corpus.
45
+
46
+ ## Model details
47
+
48
+ | | |
49
+ |---|---|
50
+ | **Version / git tag** | `daydream-chess-nanogpt-micro-1` (research run `micro-r1`) |
51
+ | **Architecture** | modern char-level (RoPE, RMSNorm, bias-free) on the shared `core` engine |
52
+ | **Size** | 4 layers Β· 4 heads Β· 128 embedding dim Β· 128 context Β· dropout 0.1 Β· **~0.79M params** |
53
+ | **Tokenizer** | character-level, **15-char** vocabulary over UCI move text on a 5Γ—5 board (files a–e, ranks 1–5, promotion letters n/q/r, space, newline) |
54
+ | **Checkpoint** | `projects/daydream/models/daydream-chess-nanogpt-micro-1/` (weights not committed) |
55
+ | **Built on** | the monorepo's shared [`core`](https://github.com/romellogoodman/sup-computer/tree/main/core) engine |
56
+ | **Developed with** | Claude ([Claude Code](https://claude.com/claude-code)) |
57
+ | **License** | MIT |
58
+
59
+ ## Intended use
60
+
61
+ Same exhibit posture as Regular, scaled to the smallest board in the
62
+ series. Pairs with `harness.py` (this folder), which plays the model
63
+ against Fairy-Stockfish under the built-in `gardner` variant.
64
+
65
+ **Out of scope.** Not a chess engine, not evaluated for playing strength.
66
+ Vocabulary and board are Gardner-minichess-specific β€” moves here are
67
+ meaningless on Regular's or Grand's boards and vice versa (see
68
+ [ADR-0022](https://github.com/romellogoodman/sup-computer/blob/main/docs/adr/0022-daydream-three-tier-sampler-prober-shape.md)
69
+ on why tiers never share a vocabulary).
70
+
71
+ ## Training data
72
+
73
+ No human corpus exists for 5Γ—5 chess, so this tier is entirely synthetic:
74
+ **4,135 self-play games** between two Fairy-Stockfish instances under the
75
+ engine's built-in `gardner` variant (bounded-depth search, not
76
+ strength-reduced β€” see
77
+ [ADR-0021](https://github.com/romellogoodman/sup-computer/blob/main/docs/adr/0021-daydream-fairy-stockfish-dependency.md)),
78
+ with randomized opening plies for game-to-game diversity (fixed-depth
79
+ search alone is fully deterministic and produced identical games on the
80
+ first attempt β€” fixed by sourcing random legal openings from the engine's
81
+ own `go perft 1` move list) plus a repetition-window cutoff for games that
82
+ fell into shuffling loops. Corpus is vendored in-folder as `games.txt`
83
+ (synthetic, seeded, code-owned β€” committed, same treatment as
84
+ kenosha-kid's `raw.txt`).
85
+
86
+ ## Training procedure
87
+
88
+ - **Optimizer:** AdamW, LR 3e-4 with cosine decay to 3e-5, 100 warmup iters, Ξ²β‚‚ 0.99, batch size 64.
89
+ - **Run:** 2,500 iterations, best val loss **0.718**.
90
+ - **Hardware:** Apple Silicon Mac (MPS / Metal backend), `torch.compile` disabled.
91
+
92
+ ## Evaluation
93
+
94
+ | Metric | Result (30 games) |
95
+ |---|---|
96
+ | **Clean completion rate** | 30/30 (100%) |
97
+ | **Legal-move rate (first try)** | 121/309 (39.2%) |
98
+
99
+ Micro's legal-move rate (39.2%) is somewhat higher than Regular's (35.3%,
100
+ [`daydream-chess-nanogpt-1`](https://github.com/romellogoodman/sup-computer/blob/main/research-docs/model-cards/daydream-chess-nanogpt-1.md)) β€” consistent
101
+ with a smaller board and smaller per-position legal-move count being an
102
+ easier legality-learning problem, though the two aren't a strict
103
+ apples-to-apples comparison (different corpora, different vocab sizes,
104
+ different training run lengths).
105
+
106
+ ## Limitations
107
+
108
+ - **Not evaluated for playing strength**, deliberately.
109
+ - **Synthetic corpus only** β€” no human Gardner-minichess games exist to
110
+ compare against; the training distribution is entirely a product of
111
+ bounded-depth Fairy-Stockfish self-play plus randomized openings.
112
+ - **Legality is learned, not guaranteed** β€” same resample-then-force-random
113
+ fallback as every tier in this series.
114
+ - **No weights in the tree** ([ADR-0002](https://github.com/romellogoodman/sup-computer/blob/main/docs/adr/0002-no-weights-in-tree.md)).
115
+
116
+ ## How to reproduce
117
+
118
+ ```bash
119
+ cd projects/daydream/models/daydream-chess-nanogpt-micro-1
120
+ python prepare.py # -> micro/{train,val}.bin + meta.pkl
121
+ python train.py config.py # -> ./ckpt.pt (2500 iters, val ~0.72)
122
+ python harness.py --games 30 # verification
123
+ ```
124
+
125
+ Requires Fairy-Stockfish on `PATH` (`brew install fairy-stockfish`).
126
+
127
+ Experiment write-up: [Can a chess model's illegal moves be the point?](https://github.com/romellogoodman/sup-computer/blob/main/research-docs/reports/illegal-moves-are-the-point.md)
128
+
129
+ ## Citation / credits
130
+
131
+ - The shared `core` engine (modern nanoGPT lineage β€” RoPE, RMSNorm, bias-free).
132
+ - [Fairy-Stockfish](https://github.com/fairy-stockfish/Fairy-Stockfish) β€” self-play corpus generator and legality arbiter, via its built-in `gardner` variant.
133
+ - Set up and trained with Claude ([Claude Code](https://claude.com/claude-code)).
ckpt.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9c1eea5e6937d378e6b1d86563e374e3314df4069ab78605982e598fa30b6860
3
+ size 9508438
daydream-chess-nanogpt-micro-1.int8.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:007f1f07eb2682010ac272af73f20b7d64b2eeef82687155114fed47497b68ce
3
+ size 1019493
daydream-chess-nanogpt-micro-1.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1f67db1e11dedb23d80220e3dce46fffd4fd6bafdf7cef57eee87b19d333b1b1
3
+ size 3312555
daydream-chess-nanogpt-micro-1.vocab.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"stoi": {"\n": 0, " ": 1, "1": 2, "2": 3, "3": 4, "4": 5, "5": 6, "a": 7, "b": 8, "c": 9, "d": 10, "e": 11, "n": 12, "q": 13, "r": 14}, "itos": {"0": "\n", "1": " ", "2": "1", "3": "2", "4": "3", "5": "4", "6": "5", "7": "a", "8": "b", "9": "c", "10": "d", "11": "e", "12": "n", "13": "q", "14": "r"}}