shubhxho commited on
Commit
a98b020
Β·
verified Β·
1 Parent(s): 6fd126b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +155 -21
README.md CHANGED
@@ -1,13 +1,53 @@
1
  ---
2
  license: mit
 
 
 
 
 
3
  tags:
4
  - chess
 
5
  - nnue
6
  - mlx
 
7
  - quantization
8
  - int8
9
  - distillation
10
- library_name: mlx
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
11
  ---
12
 
13
  # Sable β€” a 60 KB standalone chess evaluation network
@@ -32,11 +72,53 @@ constant multiplying the output layer β€” which cannot change which position the
32
  network prefers β€” was worth about 60 Elo, and getting it wrong had been
33
  poisoning every architecture comparison in this project for months.
34
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
35
  ## The input set is the whole story
36
 
37
  The obvious design uses the standard NNUE input: 768 binary features, one per
38
- (piece, colour, square). Built that way, at this size, the network plays **165
39
- Elo worse** than the hand-crafted evaluator it replaces.
 
40
 
41
  The instinct is that it's too small. It isn't. Sweeping the hidden layer from 16
42
  to 128 neurons β€” an 8x range β€” barely moves the fit against the teacher; r sits
@@ -50,15 +132,18 @@ neurons beat 32 by +12.5 Elo [+0.1, +24.9] over 3000 games and 128 beat 64 by
50
  nothing at all. The plateau is real; it starts one doubling later than this
51
  sweep said, and the fit numbers never showed the difference.
52
 
53
- Piece-square features describe where pieces **are**. Almost everything that
54
- decides a chess position is about where they can **go**. A knight's value swings
55
- wildly with what it attacks; a rook's with whether its file is open. Neither is
56
- recoverable from a one-hot square index at any width.
 
 
57
 
58
- So the budget went into the input. Alongside the 768 piece-square planes sit 166
59
- rows encoding mobility, passed pawns by rank, isolated and doubled pawns, rooks
60
- on open and half-open files, the bishop pair, king attackers and king shelter.
61
- Each row costs 64 bytes.
 
62
 
63
  | Input set | Size | r vs teacher | MAE | RMSE |
64
  |---|---|---|---|---|
@@ -261,11 +346,16 @@ against a better opponent is the expected shape of an improvement.
261
 
262
  ### Features come from the engine, never from the trainer
263
 
264
- The trainer does not compute features. It asks the engine for them through a
265
- `featdump` command that emits the active indices per position. Two
266
- implementations of one feature map is a bug class that yields a network which
267
- loads, runs, and is quietly wrong β€” very hard to find afterwards. `src/net.rs`
268
- is the single source of truth for both training and inference.
 
 
 
 
 
269
 
270
  ### Quantisation-aware by construction
271
 
@@ -277,11 +367,29 @@ model.ft = mx.clip(model.ft, -127.0 / QA, 127.0 / QA)
277
  model.out = mx.clip(model.out, -127.0 / QB, 127.0 / QB)
278
  ```
279
 
280
- So the exported network computes the function the trainer converged to. Verified,
281
- not asserted: `net.bin` is replayed through an independent NumPy reference that
282
- reproduces the Rust inference operation for operation, and the two agree on
283
- 80/80 test positions. The only disagreement ever seen was Python's floor
284
- division against Rust's truncation on negative scores β€” a bug in the reference.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
285
 
286
  ## Format
287
 
@@ -375,6 +483,32 @@ the current engine.
375
  perspectives in registers for the whole feature list rather than storing the
376
  accumulator back to memory once per row.
377
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
378
  ## License
379
 
380
  MIT.
 
1
  ---
2
  license: mit
3
+ library_name: mlx
4
+ pipeline_tag: other
5
+ inference: false
6
+ language:
7
+ - en
8
  tags:
9
  - chess
10
+ - chess-engine
11
  - nnue
12
  - mlx
13
+ - apple-silicon
14
  - quantization
15
  - int8
16
  - distillation
17
+ - knowledge-distillation
18
+ - self-play
19
+ - uci
20
+ - rust
21
+ - no-std
22
+ - edge
23
+ metrics:
24
+ - elo
25
+ - pearsonr
26
+ co2_eq_emissions:
27
+ emissions: 1.4
28
+ source: "estimated: 6 minutes of Apple M-series GPU at roughly 20W, at 0.7 kgCO2eq/kWh"
29
+ training_type: "distillation from self-play search labels"
30
+ geographical_location: "India"
31
+ hardware_used: "Apple M-series (MLX, unified memory)"
32
+ model-index:
33
+ - name: sable-chess-net
34
+ results:
35
+ - task:
36
+ type: other
37
+ name: Chess play (engine strength)
38
+ dataset:
39
+ type: self-play
40
+ name: Sable self-play, 10.1M deduplicated positions
41
+ metrics:
42
+ - type: elo
43
+ name: Elo vs previous release (3000 games, 20k nodes/move)
44
+ value: 63.0
45
+ - type: elo
46
+ name: Elo anchored to Stockfish UCI_Elo (200 games per setting, 100ms/move)
47
+ value: 2800
48
+ - type: pearsonr
49
+ name: Correlation with teacher search (rank quality, gain-invariant)
50
+ value: 0.9803
51
  ---
52
 
53
  # Sable β€” a 60 KB standalone chess evaluation network
 
72
  network prefers β€” was worth about 60 Elo, and getting it wrong had been
73
  poisoning every architecture comparison in this project for months.
74
 
75
+ ## What this is, in one paragraph
76
+
77
+ It's the evaluation function out of a chess engine I wrote. The engine searched
78
+ its own games, and this network was trained to guess what that search would have
79
+ said without doing the search. It's 60 KB of int8 weights, it runs on integer
80
+ SIMD with no framework underneath it, and it is not a PyTorch model β€” you can't
81
+ `from_pretrained` it. If you want to *use* it, you want
82
+ [the engine](https://github.com/shubhxho/sable); if you want to *read* it, the
83
+ [format section](#format) is complete enough to parse `net.bin` in twenty lines
84
+ of NumPy, which is included below.
85
+
86
+ - **Developed by:** [@shubhxho](https://huggingface.co/shubhxho)
87
+ - **Model type:** quantised int8 feedforward evaluation network (NNUE-style), distilled from tree search
88
+ - **Inputs:** 934 sparse binary features per side, computed by the engine
89
+ - **Output:** one scalar, centipawns, from the side to move's point of view
90
+ - **Trained with:** [MLX](https://github.com/ml-explore/mlx) on Apple silicon
91
+ - **License:** MIT
92
+ - **Repository:** https://github.com/shubhxho/sable
93
+
94
+ ## What it's for, and what it isn't
95
+
96
+ **Use it for:** running the Sable engine; reading a small, complete, honestly
97
+ documented example of a quantisation-aware distilled evaluation; lifting the
98
+ format or the training loop for your own engine. The whole thing is MIT and I'd
99
+ be glad to see it reused.
100
+
101
+ **Don't expect it to:** work as a general chess model, produce moves on its own,
102
+ or load into a transformers pipeline. It has no notion of a legal move. Hand it
103
+ a position and it returns a number; everything that makes that number useful β€”
104
+ move generation, search, pruning, time management β€” lives in the engine, and the
105
+ number is close to meaningless without it. The gain section below is a long
106
+ argument for exactly that point: the same weights are worth a hundred Elo more
107
+ or less depending on the search wrapped around them.
108
+
109
+ **Bias and risk, honestly:** it's a chess evaluator. The realistic harm is
110
+ someone cheating at online chess with it, which is true of every engine ever
111
+ published and which this one is far too weak to be attractive for. The more
112
+ interesting caveat is epistemic: it was distilled entirely from its own search,
113
+ so it has inherited that search's blind spots and there is no external teacher
114
+ anywhere in the loop to catch them.
115
+
116
  ## The input set is the whole story
117
 
118
  The obvious design uses the standard NNUE input: 768 binary features, one per
119
+ (piece, colour, square). I built that first. At this size it played **165 Elo
120
+ worse** than the hand-crafted evaluator it was supposed to replace, which was a
121
+ memorable afternoon.
122
 
123
  The instinct is that it's too small. It isn't. Sweeping the hidden layer from 16
124
  to 128 neurons β€” an 8x range β€” barely moves the fit against the teacher; r sits
 
132
  nothing at all. The plateau is real; it starts one doubling later than this
133
  sweep said, and the fit numbers never showed the difference.
134
 
135
+ Here's the actual problem. Piece-square features describe where pieces **are**,
136
+ and almost everything that decides a chess position is about where they can
137
+ **go**. A knight on d5 is worth wildly different amounts depending on what it
138
+ attacks. A rook is worth much more on an open file. Neither fact is recoverable
139
+ from a one-hot square index, no matter how wide you make the layer behind it β€”
140
+ the information simply isn't in the input.
141
 
142
+ So the budget went into the input instead of the hidden layer. Alongside the 768
143
+ piece-square planes sit 166 rows encoding mobility, passed pawns by rank,
144
+ isolated and doubled pawns, rooks on open and half-open files, the bishop pair,
145
+ king attackers and king shelter β€” all computed from the board by the engine and
146
+ looked up in the same embedding table. Each row costs 64 bytes.
147
 
148
  | Input set | Size | r vs teacher | MAE | RMSE |
149
  |---|---|---|---|---|
 
346
 
347
  ### Features come from the engine, never from the trainer
348
 
349
+ The trainer doesn't compute features. It asks the engine for them, through a
350
+ `featdump` command that dumps the active indices for each position, and reads
351
+ them back.
352
+
353
+ This is worth the awkwardness. Two implementations of one feature map is a bug
354
+ class where the trainer and the engine quietly disagree about what feature 431
355
+ means, and what you get is a network that loads cleanly, runs at full speed, and
356
+ plays slightly badly for reasons nothing will point you at. I would rather pipe
357
+ a gigabyte of indices through a subprocess than debug that. `src/net.rs` is the
358
+ single source of truth for both sides.
359
 
360
  ### Quantisation-aware by construction
361
 
 
367
  model.out = mx.clip(model.out, -127.0 / QB, 127.0 / QB)
368
  ```
369
 
370
+ So the exported network computes the function the trainer actually converged to,
371
+ rather than a rounded-off approximation of it.
372
+
373
+ Verified rather than asserted: `net.bin` gets replayed through an independent
374
+ NumPy reference that reproduces the Rust inference operation for operation, and
375
+ the two agree on 80/80 test positions. The only disagreement that check has ever
376
+ turned up was Python's floor division against Rust's truncation on negative
377
+ scores β€” which was a bug in the reference, not the engine, and exactly the kind
378
+ of thing the check exists to find.
379
+
380
+ ## How to get started
381
+
382
+ The fastest path is the engine itself:
383
+
384
+ ```bash
385
+ git clone https://github.com/shubhxho/sable && cd sable
386
+ cargo build --release
387
+ ./target/release/sable # then speak UCI, or type `bench 13`, `eval`, `d`
388
+ ```
389
+
390
+ `net.bin` is baked into the binary with `include_bytes!`, so the build already
391
+ contains this network β€” there's nothing to download at runtime. To read the
392
+ weights directly instead, see the NumPy snippet under [Format](#format).
393
 
394
  ## Format
395
 
 
483
  perspectives in registers for the whole feature list rather than storing the
484
  accumulator back to memory once per row.
485
 
486
+ ## Environmental impact
487
+
488
+ Rounding to something honest: about **six minutes** of Apple M-series GPU time
489
+ per training run, on hardware that draws roughly 20W doing this. Call it 1.4
490
+ gCO2eq β€” a gram and a half, less than boiling a mug of water. The 10.1M-position
491
+ dataset it trains on took considerably longer to generate than the network takes
492
+ to train, and the arena matches behind the Elo figures in this card dwarf both:
493
+ several tens of thousands of games at 20,000 nodes each. If you want the real
494
+ carbon cost of this project, it's in the measurement, not the training.
495
+
496
+ ## Citation
497
+
498
+ ```bibtex
499
+ @software{sable_chess_net,
500
+ author = {shubhxho},
501
+ title = {Sable: a 60 KB distilled chess evaluation network},
502
+ year = {2026},
503
+ url = {https://github.com/shubhxho/sable},
504
+ note = {Trained with MLX on Apple silicon; int8 quantisation-aware}
505
+ }
506
+ ```
507
+
508
+ ## Contact
509
+
510
+ Issues and questions: https://github.com/shubhxho/sable/issues
511
+
512
  ## License
513
 
514
  MIT.