shubhxho commited on
Commit
bb8a356
·
verified ·
1 Parent(s): 20b258b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +114 -28
README.md CHANGED
@@ -10,19 +10,28 @@ tags:
10
  library_name: mlx
11
  ---
12
 
13
- # Sable — a 30 KB standalone chess evaluation network
14
 
15
  The complete evaluation function for the [Sable](https://github.com/shubhxho/sable)
16
  chess engine, distilled from that engine's own search and trained with
17
  [MLX](https://github.com/ml-explore/mlx) on Apple silicon.
18
 
19
- **30,512 bytes.** It is the entire evaluation — there is no hand-crafted term
20
  underneath it, and no framework needed to run it.
21
 
22
  ```
23
- 934 features -> 32 hidden (per perspective, shared) -> clipped ReLU -> 1 of 8 output buckets
24
  ```
25
 
 
 
 
 
 
 
 
 
 
26
  ## The input set is the whole story
27
 
28
  The obvious design uses the standard NNUE input: 768 binary features, one per
@@ -31,8 +40,15 @@ Elo worse** than the hand-crafted evaluator it replaces.
31
 
32
  The instinct is that it's too small. It isn't. Sweeping the hidden layer from 16
33
  to 128 neurons — an 8x range — barely moves the fit against the teacher; r sits
34
- near 0.93 the whole way. That flatness is the finding: capacity was never the
35
- constraint.
 
 
 
 
 
 
 
36
 
37
  Piece-square features describe where pieces **are**. Almost everything that
38
  decides a chess position is about where they can **go**. A knight's value swings
@@ -42,7 +58,7 @@ recoverable from a one-hot square index at any width.
42
  So the budget went into the input. Alongside the 768 piece-square planes sit 166
43
  rows encoding mobility, passed pawns by rank, isolated and doubled pawns, rooks
44
  on open and half-open files, the bishop pair, king attackers and king shelter.
45
- Each row costs 32 bytes.
46
 
47
  | Input set | Size | r vs teacher | MAE | RMSE |
48
  |---|---|---|---|---|
@@ -64,9 +80,13 @@ load, from randomised openings, colours swapped on every pair:
64
  | 768-feature net **correcting** hand-crafted eval | +57 ± 28 Elo (600 games) |
65
  | **934-feature standalone net** vs hand-crafted eval | +35 ± 34 Elo (400 games) |
66
  | **934-feature standalone** vs the 768-feature hybrid | −3 ± 34 Elo (400 games) |
 
 
 
 
67
 
68
- The last row is the one that decided what ships. The standalone network is
69
- statistically indistinguishable from the hybrid in games, while carrying no
70
  hand-crafted evaluation at all and tracking the teacher considerably better. The
71
  same feature idea that turned a −165 Elo replacement into a viable one is what
72
  makes the standalone version possible.
@@ -76,6 +96,56 @@ Worth being straight about: the standalone net fits the teacher much better
76
  against a search's output is not the same thing as better move ordering inside
77
  one, and these match lengths cannot resolve a difference this small.
78
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
79
  ## Architecture
80
 
81
  - **Perspective pairing**: features are built twice per position, once from each
@@ -91,12 +161,12 @@ one, and these match lengths cannot resolve a difference this small.
91
 
92
  | Tensor | Shape | Type | Bytes |
93
  |---|---|---|---|
94
- | `ft_w` | 934 x 32 | int8 | 29,888 |
95
- | `ft_b` | 32 | int16 | 64 |
96
- | `out_w` | 8 x 64 | int8 | 512 |
97
  | `out_b` | 8 | int32 | 32 |
98
  | header | magic, inputs, hidden, buckets | uint32 | 16 |
99
- | | | **total** | **30,512** |
100
 
101
  ### Feature-space layout
102
 
@@ -143,8 +213,15 @@ fits in L1 cache rather than a TPU pod. The student never searches.
143
  - **Objective**: MSE in win-probability space,
144
  `sigmoid(net / 400)` against `0.9 * sigmoid(search / 400) + 0.1 * result`.
145
  - **Optimiser**: AdamW, batch 16384, lr 1e-2 with one warmup epoch then cosine
146
- decay over 30 epochs. 5% of positions are held out; the exported network is
147
  the epoch that did best on them, not the last one.
 
 
 
 
 
 
 
148
 
149
  Data volume is not the constraint either: retraining on the full 3.36M against
150
  2M moves the fit by nothing worth reporting (r 0.970 -> 0.968, RMSE 130.1 ->
@@ -213,16 +290,21 @@ Little-endian, tightly packed, no framework dependency:
213
  ```
214
  magic u32 0x334C4253 ("SBL3")
215
  inputs u32 934
216
- hidden u32 32
217
  buckets u32 8
218
- ft_w i8[934 * 32] row-major [feature][neuron]
219
- ft_b i16[32]
220
- out_w i8[8 * 64] row-major [bucket][neuron];
221
- within a bucket, first 32 = side to move,
222
- last 32 = opponent
223
  out_b i32[8]
224
  ```
225
 
 
 
 
 
 
226
  ```python
227
  import struct, numpy as np
228
  b = open("net.bin", "rb").read()
@@ -247,7 +329,8 @@ for i in $(seq 1 9); do
247
  ./target/release/sable <<< "datagen 400000 5000 $((i*7919))" > data/shard$i.txt &
248
  done; wait
249
 
250
- python train.py 3400000 16 # dumps features via the engine, writes net.bin
 
251
  cargo build --release # net.bin is include_bytes!'d into the binary
252
  cp target/release/sable sable-std
253
 
@@ -267,12 +350,15 @@ the current engine.
267
 
268
  ## Limitations
269
 
270
- - Distilled from itself. With no external engine available, the ceiling is the
271
- teacher's own search quality rather than a stronger reference.
272
- - Computing mobility and king-attacker features used to cost real throughput:
273
- 2.5 Mnps against 3.1 for the hand-crafted evaluator on the same core. That gap
274
- is now closed and slightly reversed — `bench 13`, median of five runs, is
275
- **3.37 Mnps with the network against 3.25 without it**. A direct-mapped cache
 
 
 
276
  of finished evaluations did most of it: the search asks about the same
277
  position often enough (transpositions, re-searches, null-move verification)
278
  that a good deal of the feature extraction was repeat work. The rest came from
@@ -281,8 +367,8 @@ the current engine.
281
  a claim that a network is cheaper than a hand-crafted evaluator; it is that
282
  the cache and the extraction rewrite between them now more than cover the
283
  difference.
284
- - Accumulators are refreshed in full rather than updated incrementally. At 32
285
- neurons a matrix row is four NEON registers, and most of the 166 non-piece-
286
  square rows change on almost every move anyway, so an incremental update would
287
  only cover the piece-square part. The eval cache took the easy half of that win
288
  for a fraction of the complexity, and the refresh itself now keeps both
 
10
  library_name: mlx
11
  ---
12
 
13
+ # Sable — a 60 KB standalone chess evaluation network
14
 
15
  The complete evaluation function for the [Sable](https://github.com/shubhxho/sable)
16
  chess engine, distilled from that engine's own search and trained with
17
  [MLX](https://github.com/ml-explore/mlx) on Apple silicon.
18
 
19
+ **60,976 bytes.** It is the entire evaluation — there is no hand-crafted term
20
  underneath it, and no framework needed to run it.
21
 
22
  ```
23
+ 934 features -> 64 hidden (per perspective, shared) -> clipped ReLU -> 1 of 8 output buckets
24
  ```
25
 
26
+ The engine around it plays at roughly **2800 Elo**, anchored against Stockfish's
27
+ `UCI_Elo` settings. That anchor is worth about ±70, for reasons set out under
28
+ [Playing strength](#playing-strength).
29
+
30
+ If you read one section, make it [the output gain](#the-output-gain). A single
31
+ constant multiplying the output layer — which cannot change which position the
32
+ network prefers — was worth about 60 Elo, and getting it wrong had been
33
+ poisoning every architecture comparison in this project for months.
34
+
35
  ## The input set is the whole story
36
 
37
  The obvious design uses the standard NNUE input: 768 binary features, one per
 
40
 
41
  The instinct is that it's too small. It isn't. Sweeping the hidden layer from 16
42
  to 128 neurons — an 8x range — barely moves the fit against the teacher; r sits
43
+ near 0.93 the whole way. That flatness is the finding: capacity was not the
44
+ binding constraint.
45
+
46
+ It was not *nothing*, either, and it took a later result to separate the two.
47
+ Every width comparison here predates the output gain below, so each one measured
48
+ a wide network against a differently-scaled narrow one. Held at a fixed gain, 64
49
+ neurons beat 32 by +12.5 Elo [+0.1, +24.9] over 3000 games and 128 beat 64 by
50
+ nothing at all. The plateau is real; it starts one doubling later than this
51
+ sweep said, and the fit numbers never showed the difference.
52
 
53
  Piece-square features describe where pieces **are**. Almost everything that
54
  decides a chess position is about where they can **go**. A knight's value swings
 
58
  So the budget went into the input. Alongside the 768 piece-square planes sit 166
59
  rows encoding mobility, passed pawns by rank, isolated and doubled pawns, rooks
60
  on open and half-open files, the bishop pair, king attackers and king shelter.
61
+ Each row costs 64 bytes.
62
 
63
  | Input set | Size | r vs teacher | MAE | RMSE |
64
  |---|---|---|---|---|
 
80
  | 768-feature net **correcting** hand-crafted eval | +57 ± 28 Elo (600 games) |
81
  | **934-feature standalone net** vs hand-crafted eval | +35 ± 34 Elo (400 games) |
82
  | **934-feature standalone** vs the 768-feature hybrid | −3 ± 34 Elo (400 games) |
83
+ | **the rescaled 32-neuron net** vs the previous release | +59.6 ± 21.9, +52.2 ± 21.8 (2000 games) |
84
+ | **this network** (64 neurons) vs that | +12.5 [+0.1, +24.9] (3000 games) |
85
+ | **this network** vs the previous release | **+63.0 ± 12.6** (3000 games) |
86
+ | the same, on the clock | +42.3 ± 24.3 at 100ms, +51.6 ± 34.4 at 300ms |
87
 
88
+ The fourth row is the one that decided the shape of this network. The standalone
89
+ version is statistically indistinguishable from the hybrid in games, while carrying no
90
  hand-crafted evaluation at all and tracking the teacher considerably better. The
91
  same feature idea that turned a −165 Elo replacement into a viable one is what
92
  makes the standalone version possible.
 
96
  against a search's output is not the same thing as better move ordering inside
97
  one, and these match lengths cannot resolve a difference this small.
98
 
99
+ For an absolute figure, the engine was played against Stockfish under
100
+ `UCI_LimitStrength`, 200 games at each setting, 100ms a move:
101
+
102
+ | Stockfish `UCI_Elo` | score | implied |
103
+ |---|---|---|
104
+ | 2200 | 0.958 | 2741 |
105
+ | 2500 | 0.853 | 2805 |
106
+ | 2800 | 0.465 | **2776** |
107
+ | 3000 | 0.335 | 2881 |
108
+
109
+ Call it **2800**. The 2800 row deserves the most weight because it is nearest
110
+ parity and extrapolates least. The four anchors disagree by 140 Elo, and that
111
+ spread is the honest precision — this locates the engine on someone else's scale
112
+ rather than rating it, and it is not a CCRL or FIDE number.
113
+
114
+ ## The output gain
115
+
116
+ This is the part worth reading even if nothing else here interests you.
117
+
118
+ A network distilled from a search learns to reproduce that search's score, and
119
+ that includes reproducing its **spread**. Measured over 20,000 positions, the
120
+ previous release evaluated with a standard deviation of 549 centipawns where its
121
+ teacher sat at 654 — it had been quietly understating every position for its
122
+ whole life. Retraining the same architecture on the same data fixed that, landing
123
+ at 642, and improved every fit statistic: r from 0.9794 to 0.9811, mean error
124
+ from 102cp to 82cp.
125
+
126
+ That better network lost by **38.0 ± 21.7 over 1000 games**.
127
+
128
+ Multiplying its output layer by a constant is the only thing that then separates
129
+ the two. It cannot reorder the network's preferences — r does not move — it only
130
+ changes how loud the evaluation is. Swept at 1000 games each against the previous
131
+ release: gain 1.00 gives -38.0, 0.90 gives +7.0, 0.80 gives +43.3, 0.70 gives
132
+ +59.6, 0.60 gives +58.6, 0.55 gives +47.9.
133
+
134
+ A hundred Elo across that curve, with the network knowing exactly the same things
135
+ at every point on it. The likely mechanism is that a search never consumes a
136
+ static evaluation alone — it compares it against margins, in centipawns, for
137
+ reverse futility, razoring, null-move verification and late-move reductions.
138
+ Those margins were tuned against an evaluation that happened to speak quietly.
139
+ Fix the network's calibration without fixing them and every threshold fires in
140
+ the wrong place.
141
+
142
+ It ships at `OUT_SCALE = 0.70`, applied to the output layer at export rather than
143
+ to the score in the engine, so this file stays the single description of what the
144
+ engine computes. Rerunning the sweep against the 64-neuron network put 0.55
145
+ through 0.80 all within noise of 0.70 across another 5000 games: the plateau is
146
+ wide and did not move with the architecture. The 60 Elo comes from not being at
147
+ 1.00, not from finding a precise value.
148
+
149
  ## Architecture
150
 
151
  - **Perspective pairing**: features are built twice per position, once from each
 
161
 
162
  | Tensor | Shape | Type | Bytes |
163
  |---|---|---|---|
164
+ | `ft_w` | 934 x 64 | int8 | 59,776 |
165
+ | `ft_b` | 64 | int16 | 128 |
166
+ | `out_w` | 8 x 128 | int8 | 1,024 |
167
  | `out_b` | 8 | int32 | 32 |
168
  | header | magic, inputs, hidden, buckets | uint32 | 16 |
169
+ | | | **total** | **60,976** |
170
 
171
  ### Feature-space layout
172
 
 
213
  - **Objective**: MSE in win-probability space,
214
  `sigmoid(net / 400)` against `0.9 * sigmoid(search / 400) + 0.1 * result`.
215
  - **Optimiser**: AdamW, batch 16384, lr 1e-2 with one warmup epoch then cosine
216
+ decay over 15 epochs. 5% of positions are held out; the exported network is
217
  the epoch that did best on them, not the last one.
218
+ - **Output gain**: the exported output layer is multiplied by `OUT_SCALE`, 0.70.
219
+ This is not part of the objective and it does not change which position the
220
+ network prefers; it only makes every evaluation quieter by a constant. A
221
+ network trained to reproduce a search's score reproduces its spread as well,
222
+ and the search plays substantially worse when handed one. The same network
223
+ exported at gain 1.00 loses 38.0 ± 21.7 to the previous release; at 0.70 it
224
+ wins by 59.6 ± 21.9. See README.md for the full sweep.
225
 
226
  Data volume is not the constraint either: retraining on the full 3.36M against
227
  2M moves the fit by nothing worth reporting (r 0.970 -> 0.968, RMSE 130.1 ->
 
290
  ```
291
  magic u32 0x334C4253 ("SBL3")
292
  inputs u32 934
293
+ hidden u32 64
294
  buckets u32 8
295
+ ft_w i8[934 * 64] row-major [feature][neuron]
296
+ ft_b i16[64]
297
+ out_w i8[8 * 128] row-major [bucket][neuron];
298
+ within a bucket, first 64 = side to move,
299
+ last 64 = opponent
300
  out_b i32[8]
301
  ```
302
 
303
+ The header carries `inputs`, `hidden` and `buckets`, so read those rather than
304
+ hardcoding them — this network was 32 hidden neurons until recently and the
305
+ loader rejects a file whose header disagrees with the build rather than
306
+ misreading it.
307
+
308
  ```python
309
  import struct, numpy as np
310
  b = open("net.bin", "rb").read()
 
329
  ./target/release/sable <<< "datagen 400000 5000 $((i*7919))" > data/shard$i.txt &
330
  done; wait
331
 
332
+ python train.py 10220706 15 # dumps features via the engine, writes net.bin
333
+ # NET_H=64 and OUT_SCALE=0.70 are the defaults
334
  cargo build --release # net.bin is include_bytes!'d into the binary
335
  cp target/release/sable sable-std
336
 
 
350
 
351
  ## Limitations
352
 
353
+ - Distilled from itself. The ceiling is the engine's own search quality rather
354
+ than a stronger reference. Stockfish appears in this repository only as a
355
+ measuring stick; nothing it plays has ever been trained on.
356
+ - The 2800 figure is an anchor, not a rating. Four Stockfish settings imply
357
+ ratings spread across 140 Elo, and Stockfish's own `UCI_Elo` calibration is
358
+ approximate and fitted at longer time controls than the 100ms used here.
359
+ - Computing mobility and king-attacker features costs throughput: the engine
360
+ runs about **3.1 Mnps** at `bench 13` on one M-series core, and widening the
361
+ hidden layer to 64 neurons cost 8% per node on its own. A direct-mapped cache
362
  of finished evaluations did most of it: the search asks about the same
363
  position often enough (transpositions, re-searches, null-move verification)
364
  that a good deal of the feature extraction was repeat work. The rest came from
 
367
  a claim that a network is cheaper than a hand-crafted evaluator; it is that
368
  the cache and the extraction rewrite between them now more than cover the
369
  difference.
370
+ - Accumulators are refreshed in full rather than updated incrementally. At 64
371
+ neurons a matrix row is eight NEON registers, and most of the 166 non-piece-
372
  square rows change on almost every move anyway, so an incremental update would
373
  only cover the piece-square part. The eval cache took the easy half of that win
374
  for a fraction of the complexity, and the refresh itself now keeps both