Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -41,7 +41,7 @@ model-index:
|
|
| 41 |
metrics:
|
| 42 |
- type: elo
|
| 43 |
name: Elo at 20,000 nodes/move
|
| 44 |
-
value:
|
| 45 |
args: 3000 games, three independent opening sets, 95% CI +/-12.6
|
| 46 |
verified: false
|
| 47 |
- type: elo
|
|
@@ -120,7 +120,7 @@ underneath it, and no framework needed to run it.
|
|
| 120 |
```
|
| 121 |
|
| 122 |
The engine around it plays at roughly **2800 Elo**, anchored against Stockfish's
|
| 123 |
-
`UCI_Elo` settings. That anchor is worth about ±
|
| 124 |
[Playing strength](#playing-strength).
|
| 125 |
|
| 126 |
If you read one section, make it [the output gain](#the-output-gain). A single
|
|
@@ -212,35 +212,70 @@ four times the hidden width.
|
|
| 212 |
|
| 213 |
## Playing strength
|
| 214 |
|
| 215 |
-
|
| 216 |
-
|
|
|
|
| 217 |
|
| 218 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 219 |
|---|---|
|
| 220 |
-
|
|
| 221 |
-
|
|
| 222 |
-
|
|
| 223 |
-
|
|
| 224 |
-
|
|
| 225 |
-
|
|
| 226 |
-
|
| 227 |
-
|
| 228 |
-
|
| 229 |
-
|
| 230 |
-
|
| 231 |
-
|
| 232 |
-
|
| 233 |
-
|
| 234 |
-
|
| 235 |
-
|
| 236 |
-
|
| 237 |
-
|
| 238 |
-
|
| 239 |
-
|
| 240 |
-
|
| 241 |
-
`
|
| 242 |
-
|
| 243 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 244 |
|---|---|---|
|
| 245 |
| 2600 | 0.772 | 2812 |
|
| 246 |
| 2700 | 0.638 | 2799 |
|
|
@@ -265,6 +300,27 @@ That is why the 0.5 crossover is the defensible number: two engines scoring 0.5
|
|
| 265 |
against each other are equal by definition, and that point doesn't depend on the
|
| 266 |
slope being correct. This locates the engine on someone else's approximate scale
|
| 267 |
rather than rating it, and it is not a CCRL or FIDE number.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 268 |
|
| 269 |
## The output gain
|
| 270 |
|
|
@@ -539,11 +595,11 @@ the current engine.
|
|
| 539 |
- Distilled from itself. The ceiling is the engine's own search quality rather
|
| 540 |
than a stronger reference. Stockfish appears in this repository only as a
|
| 541 |
measuring stick; nothing it plays has ever been trained on.
|
| 542 |
-
- The 2800 figure is an anchor, not a rating.
|
| 543 |
-
ratings spread across
|
| 544 |
approximate and fitted at longer time controls than the 100ms used here.
|
| 545 |
- Computing mobility and king-attacker features costs throughput: the engine
|
| 546 |
-
runs about **3.
|
| 547 |
hidden layer to 64 neurons cost 8% per node on its own. A direct-mapped cache
|
| 548 |
of finished evaluations did most of it: the search asks about the same
|
| 549 |
position often enough (transpositions, re-searches, null-move verification)
|
|
|
|
| 41 |
metrics:
|
| 42 |
- type: elo
|
| 43 |
name: Elo at 20,000 nodes/move
|
| 44 |
+
value: 62.6
|
| 45 |
args: 3000 games, three independent opening sets, 95% CI +/-12.6
|
| 46 |
verified: false
|
| 47 |
- type: elo
|
|
|
|
| 120 |
```
|
| 121 |
|
| 122 |
The engine around it plays at roughly **2800 Elo**, anchored against Stockfish's
|
| 123 |
+
`UCI_Elo` settings. That anchor is worth about ±40, for reasons set out under
|
| 124 |
[Playing strength](#playing-strength).
|
| 125 |
|
| 126 |
If you read one section, make it [the output gain](#the-output-gain). A single
|
|
|
|
| 212 |
|
| 213 |
## Playing strength
|
| 214 |
|
| 215 |
+
Everything here is measured from randomised openings with colours swapped on
|
| 216 |
+
every pair. Fixed node counts are the default, because they don't move with
|
| 217 |
+
machine load; where a clock is used it says so.
|
| 218 |
|
| 219 |
+
### This network against the one it replaces
|
| 220 |
+
|
| 221 |
+
| Conditions | Games | Result |
|
| 222 |
+
|---|---|---|
|
| 223 |
+
| 20,000 nodes/move, three independent opening sets | 3000 | **+62.6 ± 12.6** |
|
| 224 |
+
| 100ms/move | 800 | +42.3 ± 24.3 |
|
| 225 |
+
| 300ms/move | 400 | +51.6 ± 34.4 |
|
| 226 |
+
|
| 227 |
+
The three fixed-node sets were +62.9, +65.7 and +59.3, which agree far more
|
| 228 |
+
closely than most results in this project — that is what an effect well clear of
|
| 229 |
+
the noise floor looks like. The clock figures are lower because this network is
|
| 230 |
+
8% slower per node and a node-limited match hides that by construction; about
|
| 231 |
+
twenty of the sixty-three Elo is the harness being generous.
|
| 232 |
+
|
| 233 |
+
### Across search budgets
|
| 234 |
+
|
| 235 |
+
The output gain that produces most of that win was tuned at 20,000 nodes, so the
|
| 236 |
+
obvious worry is that it only pays there. 1000 games at each budget, same
|
| 237 |
+
opponent:
|
| 238 |
+
|
| 239 |
+
| Nodes/move | Result |
|
| 240 |
|---|---|
|
| 241 |
+
| 5,000 | +27.5 ± 21.6 |
|
| 242 |
+
| 10,000 | +55.4 ± 21.8 |
|
| 243 |
+
| 20,000 | +62.6 ± 12.6 |
|
| 244 |
+
| 50,000 | +63.2 ± 21.9 |
|
| 245 |
+
| 100,000 | +62.5 ± 21.9 |
|
| 246 |
+
| 200,000 | +55.0 ± 21.8 |
|
| 247 |
+
|
| 248 |
+
Flat from 20k out to 200k, ten times past the tuning point. What falls away is
|
| 249 |
+
the shallow end, which is the right direction: a shallower search prunes less and
|
| 250 |
+
consults the static evaluation less often.
|
| 251 |
+
|
| 252 |
+
### Against every older build
|
| 253 |
+
|
| 254 |
+
600 games each at 20,000 nodes against the historical binaries, plus a
|
| 255 |
+
400-game self-play control to check the harness. The previous-release row is the
|
| 256 |
+
pooled 3000-game result from above, not a 600-game match:
|
| 257 |
+
|
| 258 |
+
| Opponent | Result |
|
| 259 |
+
|---|---|
|
| 260 |
+
| the same binary, both sides (control) | +5.2 ± 34.1 |
|
| 261 |
+
| the previous release | +62.6 ± 12.6 |
|
| 262 |
+
| `sable-new` | +112.7 ± 29.3 |
|
| 263 |
+
| `sable-old`, `sable-std` | +130.3 ± 29.8 |
|
| 264 |
+
| `sable-net` | +132.9 ± 29.9 |
|
| 265 |
+
| `sable-net-v1` | +150.7 ± 30.5 |
|
| 266 |
+
| `sable-hce`, the hand-crafted evaluator | +156.2 ± 30.7 |
|
| 267 |
+
|
| 268 |
+
The control is the row that makes the others readable — zero sits comfortably
|
| 269 |
+
inside its interval, so colour swapping and pair ordering aren't leaking an
|
| 270 |
+
advantage. `sable-old` and `sable-std` return byte-identical scores because they
|
| 271 |
+
evaluate every position identically and therefore play identical games.
|
| 272 |
+
|
| 273 |
+
### On an outside scale
|
| 274 |
+
|
| 275 |
+
Against Stockfish under `UCI_LimitStrength`, **300 games at each of five
|
| 276 |
+
settings**, 100ms a move:
|
| 277 |
+
|
| 278 |
+
| Stockfish `UCI_Elo` | Score | Implied |
|
| 279 |
|---|---|---|
|
| 280 |
| 2600 | 0.772 | 2812 |
|
| 281 |
| 2700 | 0.638 | 2799 |
|
|
|
|
| 300 |
against each other are equal by definition, and that point doesn't depend on the
|
| 301 |
slope being correct. This locates the engine on someone else's approximate scale
|
| 302 |
rather than rating it, and it is not a CCRL or FIDE number.
|
| 303 |
+
`tests/calibrate.py` reproduces all of it.
|
| 304 |
+
|
| 305 |
+
### Older results, kept for the record
|
| 306 |
+
|
| 307 |
+
These decided the *shape* of the network and are not measurements of what ships
|
| 308 |
+
now. Match lengths were much shorter, which is why the intervals are so wide:
|
| 309 |
+
|
| 310 |
+
| Matchup | Result |
|
| 311 |
+
|---|---|
|
| 312 |
+
| 768-feature net **replacing** hand-crafted eval | −165 ± 69 Elo (200 games) |
|
| 313 |
+
| 768-feature net **correcting** hand-crafted eval | +57 ± 28 Elo (600 games) |
|
| 314 |
+
| 934-feature standalone net vs hand-crafted eval | +35 ± 34 Elo (400 games) |
|
| 315 |
+
| 934-feature standalone vs the 768-feature hybrid | −3 ± 34 Elo (400 games) |
|
| 316 |
+
| the rescaled 32-neuron net vs the previous release | +59.6 ± 21.9, +52.2 ± 21.8 (2000 games) |
|
| 317 |
+
| this 64-neuron network vs that | +12.5 [+0.1, +24.9] (3000 games) |
|
| 318 |
+
|
| 319 |
+
The fourth row is the one that decided the architecture: the standalone network
|
| 320 |
+
was statistically indistinguishable from the hybrid while carrying no
|
| 321 |
+
hand-crafted evaluation at all. Worth being straight about — it fit the teacher
|
| 322 |
+
much better (RMSE 130 vs 161 cp) without out-playing it, and 400 games could
|
| 323 |
+
never have resolved a difference that small.
|
| 324 |
|
| 325 |
## The output gain
|
| 326 |
|
|
|
|
| 595 |
- Distilled from itself. The ceiling is the engine's own search quality rather
|
| 596 |
than a stronger reference. Stockfish appears in this repository only as a
|
| 597 |
measuring stick; nothing it plays has ever been trained on.
|
| 598 |
+
- The 2800 figure is an anchor, not a rating. Five Stockfish settings imply
|
| 599 |
+
ratings spread across 96 Elo, and Stockfish's own `UCI_Elo` calibration is
|
| 600 |
approximate and fitted at longer time controls than the 100ms used here.
|
| 601 |
- Computing mobility and king-attacker features costs throughput: the engine
|
| 602 |
+
runs about **3.3 Mnps** at `bench 13` on one M-series core, and widening the
|
| 603 |
hidden layer to 64 neurons cost 8% per node on its own. A direct-mapped cache
|
| 604 |
of finished evaluations did most of it: the search asks about the same
|
| 605 |
position often enough (transpositions, re-searches, null-move verification)
|