Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -1,13 +1,53 @@
|
|
| 1 |
---
|
| 2 |
license: mit
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
tags:
|
| 4 |
- chess
|
|
|
|
| 5 |
- nnue
|
| 6 |
- mlx
|
|
|
|
| 7 |
- quantization
|
| 8 |
- int8
|
| 9 |
- distillation
|
| 10 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
---
|
| 12 |
|
| 13 |
# Sable β a 60 KB standalone chess evaluation network
|
|
@@ -32,11 +72,53 @@ constant multiplying the output layer β which cannot change which position the
|
|
| 32 |
network prefers β was worth about 60 Elo, and getting it wrong had been
|
| 33 |
poisoning every architecture comparison in this project for months.
|
| 34 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
## The input set is the whole story
|
| 36 |
|
| 37 |
The obvious design uses the standard NNUE input: 768 binary features, one per
|
| 38 |
-
(piece, colour, square).
|
| 39 |
-
|
|
|
|
| 40 |
|
| 41 |
The instinct is that it's too small. It isn't. Sweeping the hidden layer from 16
|
| 42 |
to 128 neurons β an 8x range β barely moves the fit against the teacher; r sits
|
|
@@ -50,15 +132,18 @@ neurons beat 32 by +12.5 Elo [+0.1, +24.9] over 3000 games and 128 beat 64 by
|
|
| 50 |
nothing at all. The plateau is real; it starts one doubling later than this
|
| 51 |
sweep said, and the fit numbers never showed the difference.
|
| 52 |
|
| 53 |
-
Piece-square features describe where pieces **are**
|
| 54 |
-
decides a chess position is about where they can
|
| 55 |
-
|
| 56 |
-
|
|
|
|
|
|
|
| 57 |
|
| 58 |
-
So the budget went into the input
|
| 59 |
-
rows encoding mobility, passed pawns by rank,
|
| 60 |
-
on open and half-open files, the bishop pair,
|
| 61 |
-
|
|
|
|
| 62 |
|
| 63 |
| Input set | Size | r vs teacher | MAE | RMSE |
|
| 64 |
|---|---|---|---|---|
|
|
@@ -261,11 +346,16 @@ against a better opponent is the expected shape of an improvement.
|
|
| 261 |
|
| 262 |
### Features come from the engine, never from the trainer
|
| 263 |
|
| 264 |
-
The trainer
|
| 265 |
-
`featdump` command that
|
| 266 |
-
|
| 267 |
-
|
| 268 |
-
is the
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 269 |
|
| 270 |
### Quantisation-aware by construction
|
| 271 |
|
|
@@ -277,11 +367,29 @@ model.ft = mx.clip(model.ft, -127.0 / QA, 127.0 / QA)
|
|
| 277 |
model.out = mx.clip(model.out, -127.0 / QB, 127.0 / QB)
|
| 278 |
```
|
| 279 |
|
| 280 |
-
So the exported network computes the function the trainer converged to
|
| 281 |
-
|
| 282 |
-
|
| 283 |
-
|
| 284 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 285 |
|
| 286 |
## Format
|
| 287 |
|
|
@@ -375,6 +483,32 @@ the current engine.
|
|
| 375 |
perspectives in registers for the whole feature list rather than storing the
|
| 376 |
accumulator back to memory once per row.
|
| 377 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 378 |
## License
|
| 379 |
|
| 380 |
MIT.
|
|
|
|
| 1 |
---
|
| 2 |
license: mit
|
| 3 |
+
library_name: mlx
|
| 4 |
+
pipeline_tag: other
|
| 5 |
+
inference: false
|
| 6 |
+
language:
|
| 7 |
+
- en
|
| 8 |
tags:
|
| 9 |
- chess
|
| 10 |
+
- chess-engine
|
| 11 |
- nnue
|
| 12 |
- mlx
|
| 13 |
+
- apple-silicon
|
| 14 |
- quantization
|
| 15 |
- int8
|
| 16 |
- distillation
|
| 17 |
+
- knowledge-distillation
|
| 18 |
+
- self-play
|
| 19 |
+
- uci
|
| 20 |
+
- rust
|
| 21 |
+
- no-std
|
| 22 |
+
- edge
|
| 23 |
+
metrics:
|
| 24 |
+
- elo
|
| 25 |
+
- pearsonr
|
| 26 |
+
co2_eq_emissions:
|
| 27 |
+
emissions: 1.4
|
| 28 |
+
source: "estimated: 6 minutes of Apple M-series GPU at roughly 20W, at 0.7 kgCO2eq/kWh"
|
| 29 |
+
training_type: "distillation from self-play search labels"
|
| 30 |
+
geographical_location: "India"
|
| 31 |
+
hardware_used: "Apple M-series (MLX, unified memory)"
|
| 32 |
+
model-index:
|
| 33 |
+
- name: sable-chess-net
|
| 34 |
+
results:
|
| 35 |
+
- task:
|
| 36 |
+
type: other
|
| 37 |
+
name: Chess play (engine strength)
|
| 38 |
+
dataset:
|
| 39 |
+
type: self-play
|
| 40 |
+
name: Sable self-play, 10.1M deduplicated positions
|
| 41 |
+
metrics:
|
| 42 |
+
- type: elo
|
| 43 |
+
name: Elo vs previous release (3000 games, 20k nodes/move)
|
| 44 |
+
value: 63.0
|
| 45 |
+
- type: elo
|
| 46 |
+
name: Elo anchored to Stockfish UCI_Elo (200 games per setting, 100ms/move)
|
| 47 |
+
value: 2800
|
| 48 |
+
- type: pearsonr
|
| 49 |
+
name: Correlation with teacher search (rank quality, gain-invariant)
|
| 50 |
+
value: 0.9803
|
| 51 |
---
|
| 52 |
|
| 53 |
# Sable β a 60 KB standalone chess evaluation network
|
|
|
|
| 72 |
network prefers β was worth about 60 Elo, and getting it wrong had been
|
| 73 |
poisoning every architecture comparison in this project for months.
|
| 74 |
|
| 75 |
+
## What this is, in one paragraph
|
| 76 |
+
|
| 77 |
+
It's the evaluation function out of a chess engine I wrote. The engine searched
|
| 78 |
+
its own games, and this network was trained to guess what that search would have
|
| 79 |
+
said without doing the search. It's 60 KB of int8 weights, it runs on integer
|
| 80 |
+
SIMD with no framework underneath it, and it is not a PyTorch model β you can't
|
| 81 |
+
`from_pretrained` it. If you want to *use* it, you want
|
| 82 |
+
[the engine](https://github.com/shubhxho/sable); if you want to *read* it, the
|
| 83 |
+
[format section](#format) is complete enough to parse `net.bin` in twenty lines
|
| 84 |
+
of NumPy, which is included below.
|
| 85 |
+
|
| 86 |
+
- **Developed by:** [@shubhxho](https://huggingface.co/shubhxho)
|
| 87 |
+
- **Model type:** quantised int8 feedforward evaluation network (NNUE-style), distilled from tree search
|
| 88 |
+
- **Inputs:** 934 sparse binary features per side, computed by the engine
|
| 89 |
+
- **Output:** one scalar, centipawns, from the side to move's point of view
|
| 90 |
+
- **Trained with:** [MLX](https://github.com/ml-explore/mlx) on Apple silicon
|
| 91 |
+
- **License:** MIT
|
| 92 |
+
- **Repository:** https://github.com/shubhxho/sable
|
| 93 |
+
|
| 94 |
+
## What it's for, and what it isn't
|
| 95 |
+
|
| 96 |
+
**Use it for:** running the Sable engine; reading a small, complete, honestly
|
| 97 |
+
documented example of a quantisation-aware distilled evaluation; lifting the
|
| 98 |
+
format or the training loop for your own engine. The whole thing is MIT and I'd
|
| 99 |
+
be glad to see it reused.
|
| 100 |
+
|
| 101 |
+
**Don't expect it to:** work as a general chess model, produce moves on its own,
|
| 102 |
+
or load into a transformers pipeline. It has no notion of a legal move. Hand it
|
| 103 |
+
a position and it returns a number; everything that makes that number useful β
|
| 104 |
+
move generation, search, pruning, time management β lives in the engine, and the
|
| 105 |
+
number is close to meaningless without it. The gain section below is a long
|
| 106 |
+
argument for exactly that point: the same weights are worth a hundred Elo more
|
| 107 |
+
or less depending on the search wrapped around them.
|
| 108 |
+
|
| 109 |
+
**Bias and risk, honestly:** it's a chess evaluator. The realistic harm is
|
| 110 |
+
someone cheating at online chess with it, which is true of every engine ever
|
| 111 |
+
published and which this one is far too weak to be attractive for. The more
|
| 112 |
+
interesting caveat is epistemic: it was distilled entirely from its own search,
|
| 113 |
+
so it has inherited that search's blind spots and there is no external teacher
|
| 114 |
+
anywhere in the loop to catch them.
|
| 115 |
+
|
| 116 |
## The input set is the whole story
|
| 117 |
|
| 118 |
The obvious design uses the standard NNUE input: 768 binary features, one per
|
| 119 |
+
(piece, colour, square). I built that first. At this size it played **165 Elo
|
| 120 |
+
worse** than the hand-crafted evaluator it was supposed to replace, which was a
|
| 121 |
+
memorable afternoon.
|
| 122 |
|
| 123 |
The instinct is that it's too small. It isn't. Sweeping the hidden layer from 16
|
| 124 |
to 128 neurons β an 8x range β barely moves the fit against the teacher; r sits
|
|
|
|
| 132 |
nothing at all. The plateau is real; it starts one doubling later than this
|
| 133 |
sweep said, and the fit numbers never showed the difference.
|
| 134 |
|
| 135 |
+
Here's the actual problem. Piece-square features describe where pieces **are**,
|
| 136 |
+
and almost everything that decides a chess position is about where they can
|
| 137 |
+
**go**. A knight on d5 is worth wildly different amounts depending on what it
|
| 138 |
+
attacks. A rook is worth much more on an open file. Neither fact is recoverable
|
| 139 |
+
from a one-hot square index, no matter how wide you make the layer behind it β
|
| 140 |
+
the information simply isn't in the input.
|
| 141 |
|
| 142 |
+
So the budget went into the input instead of the hidden layer. Alongside the 768
|
| 143 |
+
piece-square planes sit 166 rows encoding mobility, passed pawns by rank,
|
| 144 |
+
isolated and doubled pawns, rooks on open and half-open files, the bishop pair,
|
| 145 |
+
king attackers and king shelter β all computed from the board by the engine and
|
| 146 |
+
looked up in the same embedding table. Each row costs 64 bytes.
|
| 147 |
|
| 148 |
| Input set | Size | r vs teacher | MAE | RMSE |
|
| 149 |
|---|---|---|---|---|
|
|
|
|
| 346 |
|
| 347 |
### Features come from the engine, never from the trainer
|
| 348 |
|
| 349 |
+
The trainer doesn't compute features. It asks the engine for them, through a
|
| 350 |
+
`featdump` command that dumps the active indices for each position, and reads
|
| 351 |
+
them back.
|
| 352 |
+
|
| 353 |
+
This is worth the awkwardness. Two implementations of one feature map is a bug
|
| 354 |
+
class where the trainer and the engine quietly disagree about what feature 431
|
| 355 |
+
means, and what you get is a network that loads cleanly, runs at full speed, and
|
| 356 |
+
plays slightly badly for reasons nothing will point you at. I would rather pipe
|
| 357 |
+
a gigabyte of indices through a subprocess than debug that. `src/net.rs` is the
|
| 358 |
+
single source of truth for both sides.
|
| 359 |
|
| 360 |
### Quantisation-aware by construction
|
| 361 |
|
|
|
|
| 367 |
model.out = mx.clip(model.out, -127.0 / QB, 127.0 / QB)
|
| 368 |
```
|
| 369 |
|
| 370 |
+
So the exported network computes the function the trainer actually converged to,
|
| 371 |
+
rather than a rounded-off approximation of it.
|
| 372 |
+
|
| 373 |
+
Verified rather than asserted: `net.bin` gets replayed through an independent
|
| 374 |
+
NumPy reference that reproduces the Rust inference operation for operation, and
|
| 375 |
+
the two agree on 80/80 test positions. The only disagreement that check has ever
|
| 376 |
+
turned up was Python's floor division against Rust's truncation on negative
|
| 377 |
+
scores β which was a bug in the reference, not the engine, and exactly the kind
|
| 378 |
+
of thing the check exists to find.
|
| 379 |
+
|
| 380 |
+
## How to get started
|
| 381 |
+
|
| 382 |
+
The fastest path is the engine itself:
|
| 383 |
+
|
| 384 |
+
```bash
|
| 385 |
+
git clone https://github.com/shubhxho/sable && cd sable
|
| 386 |
+
cargo build --release
|
| 387 |
+
./target/release/sable # then speak UCI, or type `bench 13`, `eval`, `d`
|
| 388 |
+
```
|
| 389 |
+
|
| 390 |
+
`net.bin` is baked into the binary with `include_bytes!`, so the build already
|
| 391 |
+
contains this network β there's nothing to download at runtime. To read the
|
| 392 |
+
weights directly instead, see the NumPy snippet under [Format](#format).
|
| 393 |
|
| 394 |
## Format
|
| 395 |
|
|
|
|
| 483 |
perspectives in registers for the whole feature list rather than storing the
|
| 484 |
accumulator back to memory once per row.
|
| 485 |
|
| 486 |
+
## Environmental impact
|
| 487 |
+
|
| 488 |
+
Rounding to something honest: about **six minutes** of Apple M-series GPU time
|
| 489 |
+
per training run, on hardware that draws roughly 20W doing this. Call it 1.4
|
| 490 |
+
gCO2eq β a gram and a half, less than boiling a mug of water. The 10.1M-position
|
| 491 |
+
dataset it trains on took considerably longer to generate than the network takes
|
| 492 |
+
to train, and the arena matches behind the Elo figures in this card dwarf both:
|
| 493 |
+
several tens of thousands of games at 20,000 nodes each. If you want the real
|
| 494 |
+
carbon cost of this project, it's in the measurement, not the training.
|
| 495 |
+
|
| 496 |
+
## Citation
|
| 497 |
+
|
| 498 |
+
```bibtex
|
| 499 |
+
@software{sable_chess_net,
|
| 500 |
+
author = {shubhxho},
|
| 501 |
+
title = {Sable: a 60 KB distilled chess evaluation network},
|
| 502 |
+
year = {2026},
|
| 503 |
+
url = {https://github.com/shubhxho/sable},
|
| 504 |
+
note = {Trained with MLX on Apple silicon; int8 quantisation-aware}
|
| 505 |
+
}
|
| 506 |
+
```
|
| 507 |
+
|
| 508 |
+
## Contact
|
| 509 |
+
|
| 510 |
+
Issues and questions: https://github.com/shubhxho/sable/issues
|
| 511 |
+
|
| 512 |
## License
|
| 513 |
|
| 514 |
MIT.
|