shubhxho commited on
Commit
7cdac0d
·
verified ·
1 Parent(s): 6787c51

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +88 -32
README.md CHANGED
@@ -41,7 +41,7 @@ model-index:
41
  metrics:
42
  - type: elo
43
  name: Elo at 20,000 nodes/move
44
- value: 63.0
45
  args: 3000 games, three independent opening sets, 95% CI +/-12.6
46
  verified: false
47
  - type: elo
@@ -120,7 +120,7 @@ underneath it, and no framework needed to run it.
120
  ```
121
 
122
  The engine around it plays at roughly **2800 Elo**, anchored against Stockfish's
123
- `UCI_Elo` settings. That anchor is worth about ±70, for reasons set out under
124
  [Playing strength](#playing-strength).
125
 
126
  If you read one section, make it [the output gain](#the-output-gain). A single
@@ -212,35 +212,70 @@ four times the hidden width.
212
 
213
  ## Playing strength
214
 
215
- Measured at a fixed 20,000 nodes per move so results don't move with machine
216
- load, from randomised openings, colours swapped on every pair:
 
217
 
218
- | Matchup | Result |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
219
  |---|---|
220
- | 768-feature net **replacing** hand-crafted eval | −165 ± 69 Elo (200 games) |
221
- | 768-feature net **correcting** hand-crafted eval | +57 ± 28 Elo (600 games) |
222
- | **934-feature standalone net** vs hand-crafted eval | +35 ± 34 Elo (400 games) |
223
- | **934-feature standalone** vs the 768-feature hybrid | −3 ± 34 Elo (400 games) |
224
- | **the rescaled 32-neuron net** vs the previous release | +59.6 ± 21.9, +52.2 ± 21.8 (2000 games) |
225
- | **this network** (64 neurons) vs that | +12.5 [+0.1, +24.9] (3000 games) |
226
- | **this network** vs the previous release | **+63.0 ± 12.6** (3000 games) |
227
- | the same, on the clock | +42.3 ± 24.3 at 100ms, +51.6 ± 34.4 at 300ms |
228
-
229
- The fourth row is the one that decided the shape of this network. The standalone
230
- version is statistically indistinguishable from the hybrid in games, while carrying no
231
- hand-crafted evaluation at all and tracking the teacher considerably better. The
232
- same feature idea that turned a −165 Elo replacement into a viable one is what
233
- makes the standalone version possible.
234
-
235
- Worth being straight about: the standalone net fits the teacher much better
236
- (RMSE 130 vs 161 cp) than the hybrid but does not out-play it. Better regression
237
- against a search's output is not the same thing as better move ordering inside
238
- one, and these match lengths cannot resolve a difference this small.
239
-
240
- For an absolute figure, the engine was played against Stockfish under
241
- `UCI_LimitStrength`, 200 games at each setting, 100ms a move:
242
-
243
- | Stockfish `UCI_Elo` | score | implied |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
244
  |---|---|---|
245
  | 2600 | 0.772 | 2812 |
246
  | 2700 | 0.638 | 2799 |
@@ -265,6 +300,27 @@ That is why the 0.5 crossover is the defensible number: two engines scoring 0.5
265
  against each other are equal by definition, and that point doesn't depend on the
266
  slope being correct. This locates the engine on someone else's approximate scale
267
  rather than rating it, and it is not a CCRL or FIDE number.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
268
 
269
  ## The output gain
270
 
@@ -539,11 +595,11 @@ the current engine.
539
  - Distilled from itself. The ceiling is the engine's own search quality rather
540
  than a stronger reference. Stockfish appears in this repository only as a
541
  measuring stick; nothing it plays has ever been trained on.
542
- - The 2800 figure is an anchor, not a rating. Four Stockfish settings imply
543
- ratings spread across 140 Elo, and Stockfish's own `UCI_Elo` calibration is
544
  approximate and fitted at longer time controls than the 100ms used here.
545
  - Computing mobility and king-attacker features costs throughput: the engine
546
- runs about **3.1 Mnps** at `bench 13` on one M-series core, and widening the
547
  hidden layer to 64 neurons cost 8% per node on its own. A direct-mapped cache
548
  of finished evaluations did most of it: the search asks about the same
549
  position often enough (transpositions, re-searches, null-move verification)
 
41
  metrics:
42
  - type: elo
43
  name: Elo at 20,000 nodes/move
44
+ value: 62.6
45
  args: 3000 games, three independent opening sets, 95% CI +/-12.6
46
  verified: false
47
  - type: elo
 
120
  ```
121
 
122
  The engine around it plays at roughly **2800 Elo**, anchored against Stockfish's
123
+ `UCI_Elo` settings. That anchor is worth about ±40, for reasons set out under
124
  [Playing strength](#playing-strength).
125
 
126
  If you read one section, make it [the output gain](#the-output-gain). A single
 
212
 
213
  ## Playing strength
214
 
215
+ Everything here is measured from randomised openings with colours swapped on
216
+ every pair. Fixed node counts are the default, because they don't move with
217
+ machine load; where a clock is used it says so.
218
 
219
+ ### This network against the one it replaces
220
+
221
+ | Conditions | Games | Result |
222
+ |---|---|---|
223
+ | 20,000 nodes/move, three independent opening sets | 3000 | **+62.6 ± 12.6** |
224
+ | 100ms/move | 800 | +42.3 ± 24.3 |
225
+ | 300ms/move | 400 | +51.6 ± 34.4 |
226
+
227
+ The three fixed-node sets were +62.9, +65.7 and +59.3, which agree far more
228
+ closely than most results in this project — that is what an effect well clear of
229
+ the noise floor looks like. The clock figures are lower because this network is
230
+ 8% slower per node and a node-limited match hides that by construction; about
231
+ twenty of the sixty-three Elo is the harness being generous.
232
+
233
+ ### Across search budgets
234
+
235
+ The output gain that produces most of that win was tuned at 20,000 nodes, so the
236
+ obvious worry is that it only pays there. 1000 games at each budget, same
237
+ opponent:
238
+
239
+ | Nodes/move | Result |
240
  |---|---|
241
+ | 5,000 | +27.5 ± 21.6 |
242
+ | 10,000 | +55.4 ± 21.8 |
243
+ | 20,000 | +62.6 ± 12.6 |
244
+ | 50,000 | +63.2 ± 21.9 |
245
+ | 100,000 | +62.5 ± 21.9 |
246
+ | 200,000 | +55.0 ± 21.8 |
247
+
248
+ Flat from 20k out to 200k, ten times past the tuning point. What falls away is
249
+ the shallow end, which is the right direction: a shallower search prunes less and
250
+ consults the static evaluation less often.
251
+
252
+ ### Against every older build
253
+
254
+ 600 games each at 20,000 nodes against the historical binaries, plus a
255
+ 400-game self-play control to check the harness. The previous-release row is the
256
+ pooled 3000-game result from above, not a 600-game match:
257
+
258
+ | Opponent | Result |
259
+ |---|---|
260
+ | the same binary, both sides (control) | +5.2 ± 34.1 |
261
+ | the previous release | +62.6 ± 12.6 |
262
+ | `sable-new` | +112.7 ± 29.3 |
263
+ | `sable-old`, `sable-std` | +130.3 ± 29.8 |
264
+ | `sable-net` | +132.9 ± 29.9 |
265
+ | `sable-net-v1` | +150.7 ± 30.5 |
266
+ | `sable-hce`, the hand-crafted evaluator | +156.2 ± 30.7 |
267
+
268
+ The control is the row that makes the others readable — zero sits comfortably
269
+ inside its interval, so colour swapping and pair ordering aren't leaking an
270
+ advantage. `sable-old` and `sable-std` return byte-identical scores because they
271
+ evaluate every position identically and therefore play identical games.
272
+
273
+ ### On an outside scale
274
+
275
+ Against Stockfish under `UCI_LimitStrength`, **300 games at each of five
276
+ settings**, 100ms a move:
277
+
278
+ | Stockfish `UCI_Elo` | Score | Implied |
279
  |---|---|---|
280
  | 2600 | 0.772 | 2812 |
281
  | 2700 | 0.638 | 2799 |
 
300
  against each other are equal by definition, and that point doesn't depend on the
301
  slope being correct. This locates the engine on someone else's approximate scale
302
  rather than rating it, and it is not a CCRL or FIDE number.
303
+ `tests/calibrate.py` reproduces all of it.
304
+
305
+ ### Older results, kept for the record
306
+
307
+ These decided the *shape* of the network and are not measurements of what ships
308
+ now. Match lengths were much shorter, which is why the intervals are so wide:
309
+
310
+ | Matchup | Result |
311
+ |---|---|
312
+ | 768-feature net **replacing** hand-crafted eval | −165 ± 69 Elo (200 games) |
313
+ | 768-feature net **correcting** hand-crafted eval | +57 ± 28 Elo (600 games) |
314
+ | 934-feature standalone net vs hand-crafted eval | +35 ± 34 Elo (400 games) |
315
+ | 934-feature standalone vs the 768-feature hybrid | −3 ± 34 Elo (400 games) |
316
+ | the rescaled 32-neuron net vs the previous release | +59.6 ± 21.9, +52.2 ± 21.8 (2000 games) |
317
+ | this 64-neuron network vs that | +12.5 [+0.1, +24.9] (3000 games) |
318
+
319
+ The fourth row is the one that decided the architecture: the standalone network
320
+ was statistically indistinguishable from the hybrid while carrying no
321
+ hand-crafted evaluation at all. Worth being straight about — it fit the teacher
322
+ much better (RMSE 130 vs 161 cp) without out-playing it, and 400 games could
323
+ never have resolved a difference that small.
324
 
325
  ## The output gain
326
 
 
595
  - Distilled from itself. The ceiling is the engine's own search quality rather
596
  than a stronger reference. Stockfish appears in this repository only as a
597
  measuring stick; nothing it plays has ever been trained on.
598
+ - The 2800 figure is an anchor, not a rating. Five Stockfish settings imply
599
+ ratings spread across 96 Elo, and Stockfish's own `UCI_Elo` calibration is
600
  approximate and fitted at longer time controls than the 100ms used here.
601
  - Computing mobility and king-attacker features costs throughput: the engine
602
+ runs about **3.3 Mnps** at `bench 13` on one M-series core, and widening the
603
  hidden layer to 64 neurons cost 8% per node on its own. A direct-mapped cache
604
  of finished evaluations did most of it: the search asks about the same
605
  position often enough (transpositions, re-searches, null-move verification)