Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -10,19 +10,28 @@ tags:
|
|
| 10 |
library_name: mlx
|
| 11 |
---
|
| 12 |
|
| 13 |
-
# Sable — a
|
| 14 |
|
| 15 |
The complete evaluation function for the [Sable](https://github.com/shubhxho/sable)
|
| 16 |
chess engine, distilled from that engine's own search and trained with
|
| 17 |
[MLX](https://github.com/ml-explore/mlx) on Apple silicon.
|
| 18 |
|
| 19 |
-
**
|
| 20 |
underneath it, and no framework needed to run it.
|
| 21 |
|
| 22 |
```
|
| 23 |
-
934 features ->
|
| 24 |
```
|
| 25 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 26 |
## The input set is the whole story
|
| 27 |
|
| 28 |
The obvious design uses the standard NNUE input: 768 binary features, one per
|
|
@@ -31,8 +40,15 @@ Elo worse** than the hand-crafted evaluator it replaces.
|
|
| 31 |
|
| 32 |
The instinct is that it's too small. It isn't. Sweeping the hidden layer from 16
|
| 33 |
to 128 neurons — an 8x range — barely moves the fit against the teacher; r sits
|
| 34 |
-
near 0.93 the whole way. That flatness is the finding: capacity was
|
| 35 |
-
constraint.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
|
| 37 |
Piece-square features describe where pieces **are**. Almost everything that
|
| 38 |
decides a chess position is about where they can **go**. A knight's value swings
|
|
@@ -42,7 +58,7 @@ recoverable from a one-hot square index at any width.
|
|
| 42 |
So the budget went into the input. Alongside the 768 piece-square planes sit 166
|
| 43 |
rows encoding mobility, passed pawns by rank, isolated and doubled pawns, rooks
|
| 44 |
on open and half-open files, the bishop pair, king attackers and king shelter.
|
| 45 |
-
Each row costs
|
| 46 |
|
| 47 |
| Input set | Size | r vs teacher | MAE | RMSE |
|
| 48 |
|---|---|---|---|---|
|
|
@@ -64,9 +80,13 @@ load, from randomised openings, colours swapped on every pair:
|
|
| 64 |
| 768-feature net **correcting** hand-crafted eval | +57 ± 28 Elo (600 games) |
|
| 65 |
| **934-feature standalone net** vs hand-crafted eval | +35 ± 34 Elo (400 games) |
|
| 66 |
| **934-feature standalone** vs the 768-feature hybrid | −3 ± 34 Elo (400 games) |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 67 |
|
| 68 |
-
The
|
| 69 |
-
statistically indistinguishable from the hybrid in games, while carrying no
|
| 70 |
hand-crafted evaluation at all and tracking the teacher considerably better. The
|
| 71 |
same feature idea that turned a −165 Elo replacement into a viable one is what
|
| 72 |
makes the standalone version possible.
|
|
@@ -76,6 +96,56 @@ Worth being straight about: the standalone net fits the teacher much better
|
|
| 76 |
against a search's output is not the same thing as better move ordering inside
|
| 77 |
one, and these match lengths cannot resolve a difference this small.
|
| 78 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 79 |
## Architecture
|
| 80 |
|
| 81 |
- **Perspective pairing**: features are built twice per position, once from each
|
|
@@ -91,12 +161,12 @@ one, and these match lengths cannot resolve a difference this small.
|
|
| 91 |
|
| 92 |
| Tensor | Shape | Type | Bytes |
|
| 93 |
|---|---|---|---|
|
| 94 |
-
| `ft_w` | 934 x
|
| 95 |
-
| `ft_b` |
|
| 96 |
-
| `out_w` | 8 x
|
| 97 |
| `out_b` | 8 | int32 | 32 |
|
| 98 |
| header | magic, inputs, hidden, buckets | uint32 | 16 |
|
| 99 |
-
| | | **total** | **
|
| 100 |
|
| 101 |
### Feature-space layout
|
| 102 |
|
|
@@ -143,8 +213,15 @@ fits in L1 cache rather than a TPU pod. The student never searches.
|
|
| 143 |
- **Objective**: MSE in win-probability space,
|
| 144 |
`sigmoid(net / 400)` against `0.9 * sigmoid(search / 400) + 0.1 * result`.
|
| 145 |
- **Optimiser**: AdamW, batch 16384, lr 1e-2 with one warmup epoch then cosine
|
| 146 |
-
decay over
|
| 147 |
the epoch that did best on them, not the last one.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 148 |
|
| 149 |
Data volume is not the constraint either: retraining on the full 3.36M against
|
| 150 |
2M moves the fit by nothing worth reporting (r 0.970 -> 0.968, RMSE 130.1 ->
|
|
@@ -213,16 +290,21 @@ Little-endian, tightly packed, no framework dependency:
|
|
| 213 |
```
|
| 214 |
magic u32 0x334C4253 ("SBL3")
|
| 215 |
inputs u32 934
|
| 216 |
-
hidden u32
|
| 217 |
buckets u32 8
|
| 218 |
-
ft_w i8[934 *
|
| 219 |
-
ft_b i16[
|
| 220 |
-
out_w i8[8 *
|
| 221 |
-
within a bucket, first
|
| 222 |
-
last
|
| 223 |
out_b i32[8]
|
| 224 |
```
|
| 225 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 226 |
```python
|
| 227 |
import struct, numpy as np
|
| 228 |
b = open("net.bin", "rb").read()
|
|
@@ -247,7 +329,8 @@ for i in $(seq 1 9); do
|
|
| 247 |
./target/release/sable <<< "datagen 400000 5000 $((i*7919))" > data/shard$i.txt &
|
| 248 |
done; wait
|
| 249 |
|
| 250 |
-
python train.py
|
|
|
|
| 251 |
cargo build --release # net.bin is include_bytes!'d into the binary
|
| 252 |
cp target/release/sable sable-std
|
| 253 |
|
|
@@ -267,12 +350,15 @@ the current engine.
|
|
| 267 |
|
| 268 |
## Limitations
|
| 269 |
|
| 270 |
-
- Distilled from itself.
|
| 271 |
-
|
| 272 |
-
|
| 273 |
-
|
| 274 |
-
|
| 275 |
-
|
|
|
|
|
|
|
|
|
|
| 276 |
of finished evaluations did most of it: the search asks about the same
|
| 277 |
position often enough (transpositions, re-searches, null-move verification)
|
| 278 |
that a good deal of the feature extraction was repeat work. The rest came from
|
|
@@ -281,8 +367,8 @@ the current engine.
|
|
| 281 |
a claim that a network is cheaper than a hand-crafted evaluator; it is that
|
| 282 |
the cache and the extraction rewrite between them now more than cover the
|
| 283 |
difference.
|
| 284 |
-
- Accumulators are refreshed in full rather than updated incrementally. At
|
| 285 |
-
neurons a matrix row is
|
| 286 |
square rows change on almost every move anyway, so an incremental update would
|
| 287 |
only cover the piece-square part. The eval cache took the easy half of that win
|
| 288 |
for a fraction of the complexity, and the refresh itself now keeps both
|
|
|
|
| 10 |
library_name: mlx
|
| 11 |
---
|
| 12 |
|
| 13 |
+
# Sable — a 60 KB standalone chess evaluation network
|
| 14 |
|
| 15 |
The complete evaluation function for the [Sable](https://github.com/shubhxho/sable)
|
| 16 |
chess engine, distilled from that engine's own search and trained with
|
| 17 |
[MLX](https://github.com/ml-explore/mlx) on Apple silicon.
|
| 18 |
|
| 19 |
+
**60,976 bytes.** It is the entire evaluation — there is no hand-crafted term
|
| 20 |
underneath it, and no framework needed to run it.
|
| 21 |
|
| 22 |
```
|
| 23 |
+
934 features -> 64 hidden (per perspective, shared) -> clipped ReLU -> 1 of 8 output buckets
|
| 24 |
```
|
| 25 |
|
| 26 |
+
The engine around it plays at roughly **2800 Elo**, anchored against Stockfish's
|
| 27 |
+
`UCI_Elo` settings. That anchor is worth about ±70, for reasons set out under
|
| 28 |
+
[Playing strength](#playing-strength).
|
| 29 |
+
|
| 30 |
+
If you read one section, make it [the output gain](#the-output-gain). A single
|
| 31 |
+
constant multiplying the output layer — which cannot change which position the
|
| 32 |
+
network prefers — was worth about 60 Elo, and getting it wrong had been
|
| 33 |
+
poisoning every architecture comparison in this project for months.
|
| 34 |
+
|
| 35 |
## The input set is the whole story
|
| 36 |
|
| 37 |
The obvious design uses the standard NNUE input: 768 binary features, one per
|
|
|
|
| 40 |
|
| 41 |
The instinct is that it's too small. It isn't. Sweeping the hidden layer from 16
|
| 42 |
to 128 neurons — an 8x range — barely moves the fit against the teacher; r sits
|
| 43 |
+
near 0.93 the whole way. That flatness is the finding: capacity was not the
|
| 44 |
+
binding constraint.
|
| 45 |
+
|
| 46 |
+
It was not *nothing*, either, and it took a later result to separate the two.
|
| 47 |
+
Every width comparison here predates the output gain below, so each one measured
|
| 48 |
+
a wide network against a differently-scaled narrow one. Held at a fixed gain, 64
|
| 49 |
+
neurons beat 32 by +12.5 Elo [+0.1, +24.9] over 3000 games and 128 beat 64 by
|
| 50 |
+
nothing at all. The plateau is real; it starts one doubling later than this
|
| 51 |
+
sweep said, and the fit numbers never showed the difference.
|
| 52 |
|
| 53 |
Piece-square features describe where pieces **are**. Almost everything that
|
| 54 |
decides a chess position is about where they can **go**. A knight's value swings
|
|
|
|
| 58 |
So the budget went into the input. Alongside the 768 piece-square planes sit 166
|
| 59 |
rows encoding mobility, passed pawns by rank, isolated and doubled pawns, rooks
|
| 60 |
on open and half-open files, the bishop pair, king attackers and king shelter.
|
| 61 |
+
Each row costs 64 bytes.
|
| 62 |
|
| 63 |
| Input set | Size | r vs teacher | MAE | RMSE |
|
| 64 |
|---|---|---|---|---|
|
|
|
|
| 80 |
| 768-feature net **correcting** hand-crafted eval | +57 ± 28 Elo (600 games) |
|
| 81 |
| **934-feature standalone net** vs hand-crafted eval | +35 ± 34 Elo (400 games) |
|
| 82 |
| **934-feature standalone** vs the 768-feature hybrid | −3 ± 34 Elo (400 games) |
|
| 83 |
+
| **the rescaled 32-neuron net** vs the previous release | +59.6 ± 21.9, +52.2 ± 21.8 (2000 games) |
|
| 84 |
+
| **this network** (64 neurons) vs that | +12.5 [+0.1, +24.9] (3000 games) |
|
| 85 |
+
| **this network** vs the previous release | **+63.0 ± 12.6** (3000 games) |
|
| 86 |
+
| the same, on the clock | +42.3 ± 24.3 at 100ms, +51.6 ± 34.4 at 300ms |
|
| 87 |
|
| 88 |
+
The fourth row is the one that decided the shape of this network. The standalone
|
| 89 |
+
version is statistically indistinguishable from the hybrid in games, while carrying no
|
| 90 |
hand-crafted evaluation at all and tracking the teacher considerably better. The
|
| 91 |
same feature idea that turned a −165 Elo replacement into a viable one is what
|
| 92 |
makes the standalone version possible.
|
|
|
|
| 96 |
against a search's output is not the same thing as better move ordering inside
|
| 97 |
one, and these match lengths cannot resolve a difference this small.
|
| 98 |
|
| 99 |
+
For an absolute figure, the engine was played against Stockfish under
|
| 100 |
+
`UCI_LimitStrength`, 200 games at each setting, 100ms a move:
|
| 101 |
+
|
| 102 |
+
| Stockfish `UCI_Elo` | score | implied |
|
| 103 |
+
|---|---|---|
|
| 104 |
+
| 2200 | 0.958 | 2741 |
|
| 105 |
+
| 2500 | 0.853 | 2805 |
|
| 106 |
+
| 2800 | 0.465 | **2776** |
|
| 107 |
+
| 3000 | 0.335 | 2881 |
|
| 108 |
+
|
| 109 |
+
Call it **2800**. The 2800 row deserves the most weight because it is nearest
|
| 110 |
+
parity and extrapolates least. The four anchors disagree by 140 Elo, and that
|
| 111 |
+
spread is the honest precision — this locates the engine on someone else's scale
|
| 112 |
+
rather than rating it, and it is not a CCRL or FIDE number.
|
| 113 |
+
|
| 114 |
+
## The output gain
|
| 115 |
+
|
| 116 |
+
This is the part worth reading even if nothing else here interests you.
|
| 117 |
+
|
| 118 |
+
A network distilled from a search learns to reproduce that search's score, and
|
| 119 |
+
that includes reproducing its **spread**. Measured over 20,000 positions, the
|
| 120 |
+
previous release evaluated with a standard deviation of 549 centipawns where its
|
| 121 |
+
teacher sat at 654 — it had been quietly understating every position for its
|
| 122 |
+
whole life. Retraining the same architecture on the same data fixed that, landing
|
| 123 |
+
at 642, and improved every fit statistic: r from 0.9794 to 0.9811, mean error
|
| 124 |
+
from 102cp to 82cp.
|
| 125 |
+
|
| 126 |
+
That better network lost by **38.0 ± 21.7 over 1000 games**.
|
| 127 |
+
|
| 128 |
+
Multiplying its output layer by a constant is the only thing that then separates
|
| 129 |
+
the two. It cannot reorder the network's preferences — r does not move — it only
|
| 130 |
+
changes how loud the evaluation is. Swept at 1000 games each against the previous
|
| 131 |
+
release: gain 1.00 gives -38.0, 0.90 gives +7.0, 0.80 gives +43.3, 0.70 gives
|
| 132 |
+
+59.6, 0.60 gives +58.6, 0.55 gives +47.9.
|
| 133 |
+
|
| 134 |
+
A hundred Elo across that curve, with the network knowing exactly the same things
|
| 135 |
+
at every point on it. The likely mechanism is that a search never consumes a
|
| 136 |
+
static evaluation alone — it compares it against margins, in centipawns, for
|
| 137 |
+
reverse futility, razoring, null-move verification and late-move reductions.
|
| 138 |
+
Those margins were tuned against an evaluation that happened to speak quietly.
|
| 139 |
+
Fix the network's calibration without fixing them and every threshold fires in
|
| 140 |
+
the wrong place.
|
| 141 |
+
|
| 142 |
+
It ships at `OUT_SCALE = 0.70`, applied to the output layer at export rather than
|
| 143 |
+
to the score in the engine, so this file stays the single description of what the
|
| 144 |
+
engine computes. Rerunning the sweep against the 64-neuron network put 0.55
|
| 145 |
+
through 0.80 all within noise of 0.70 across another 5000 games: the plateau is
|
| 146 |
+
wide and did not move with the architecture. The 60 Elo comes from not being at
|
| 147 |
+
1.00, not from finding a precise value.
|
| 148 |
+
|
| 149 |
## Architecture
|
| 150 |
|
| 151 |
- **Perspective pairing**: features are built twice per position, once from each
|
|
|
|
| 161 |
|
| 162 |
| Tensor | Shape | Type | Bytes |
|
| 163 |
|---|---|---|---|
|
| 164 |
+
| `ft_w` | 934 x 64 | int8 | 59,776 |
|
| 165 |
+
| `ft_b` | 64 | int16 | 128 |
|
| 166 |
+
| `out_w` | 8 x 128 | int8 | 1,024 |
|
| 167 |
| `out_b` | 8 | int32 | 32 |
|
| 168 |
| header | magic, inputs, hidden, buckets | uint32 | 16 |
|
| 169 |
+
| | | **total** | **60,976** |
|
| 170 |
|
| 171 |
### Feature-space layout
|
| 172 |
|
|
|
|
| 213 |
- **Objective**: MSE in win-probability space,
|
| 214 |
`sigmoid(net / 400)` against `0.9 * sigmoid(search / 400) + 0.1 * result`.
|
| 215 |
- **Optimiser**: AdamW, batch 16384, lr 1e-2 with one warmup epoch then cosine
|
| 216 |
+
decay over 15 epochs. 5% of positions are held out; the exported network is
|
| 217 |
the epoch that did best on them, not the last one.
|
| 218 |
+
- **Output gain**: the exported output layer is multiplied by `OUT_SCALE`, 0.70.
|
| 219 |
+
This is not part of the objective and it does not change which position the
|
| 220 |
+
network prefers; it only makes every evaluation quieter by a constant. A
|
| 221 |
+
network trained to reproduce a search's score reproduces its spread as well,
|
| 222 |
+
and the search plays substantially worse when handed one. The same network
|
| 223 |
+
exported at gain 1.00 loses 38.0 ± 21.7 to the previous release; at 0.70 it
|
| 224 |
+
wins by 59.6 ± 21.9. See README.md for the full sweep.
|
| 225 |
|
| 226 |
Data volume is not the constraint either: retraining on the full 3.36M against
|
| 227 |
2M moves the fit by nothing worth reporting (r 0.970 -> 0.968, RMSE 130.1 ->
|
|
|
|
| 290 |
```
|
| 291 |
magic u32 0x334C4253 ("SBL3")
|
| 292 |
inputs u32 934
|
| 293 |
+
hidden u32 64
|
| 294 |
buckets u32 8
|
| 295 |
+
ft_w i8[934 * 64] row-major [feature][neuron]
|
| 296 |
+
ft_b i16[64]
|
| 297 |
+
out_w i8[8 * 128] row-major [bucket][neuron];
|
| 298 |
+
within a bucket, first 64 = side to move,
|
| 299 |
+
last 64 = opponent
|
| 300 |
out_b i32[8]
|
| 301 |
```
|
| 302 |
|
| 303 |
+
The header carries `inputs`, `hidden` and `buckets`, so read those rather than
|
| 304 |
+
hardcoding them — this network was 32 hidden neurons until recently and the
|
| 305 |
+
loader rejects a file whose header disagrees with the build rather than
|
| 306 |
+
misreading it.
|
| 307 |
+
|
| 308 |
```python
|
| 309 |
import struct, numpy as np
|
| 310 |
b = open("net.bin", "rb").read()
|
|
|
|
| 329 |
./target/release/sable <<< "datagen 400000 5000 $((i*7919))" > data/shard$i.txt &
|
| 330 |
done; wait
|
| 331 |
|
| 332 |
+
python train.py 10220706 15 # dumps features via the engine, writes net.bin
|
| 333 |
+
# NET_H=64 and OUT_SCALE=0.70 are the defaults
|
| 334 |
cargo build --release # net.bin is include_bytes!'d into the binary
|
| 335 |
cp target/release/sable sable-std
|
| 336 |
|
|
|
|
| 350 |
|
| 351 |
## Limitations
|
| 352 |
|
| 353 |
+
- Distilled from itself. The ceiling is the engine's own search quality rather
|
| 354 |
+
than a stronger reference. Stockfish appears in this repository only as a
|
| 355 |
+
measuring stick; nothing it plays has ever been trained on.
|
| 356 |
+
- The 2800 figure is an anchor, not a rating. Four Stockfish settings imply
|
| 357 |
+
ratings spread across 140 Elo, and Stockfish's own `UCI_Elo` calibration is
|
| 358 |
+
approximate and fitted at longer time controls than the 100ms used here.
|
| 359 |
+
- Computing mobility and king-attacker features costs throughput: the engine
|
| 360 |
+
runs about **3.1 Mnps** at `bench 13` on one M-series core, and widening the
|
| 361 |
+
hidden layer to 64 neurons cost 8% per node on its own. A direct-mapped cache
|
| 362 |
of finished evaluations did most of it: the search asks about the same
|
| 363 |
position often enough (transpositions, re-searches, null-move verification)
|
| 364 |
that a good deal of the feature extraction was repeat work. The rest came from
|
|
|
|
| 367 |
a claim that a network is cheaper than a hand-crafted evaluator; it is that
|
| 368 |
the cache and the extraction rewrite between them now more than cover the
|
| 369 |
difference.
|
| 370 |
+
- Accumulators are refreshed in full rather than updated incrementally. At 64
|
| 371 |
+
neurons a matrix row is eight NEON registers, and most of the 166 non-piece-
|
| 372 |
square rows change on almost every move anyway, so an incremental update would
|
| 373 |
only cover the piece-square part. The eval cache took the easy half of that win
|
| 374 |
for a fraction of the complexity, and the refresh itself now keeps both
|