Document the +15.5 Elo network and its data
Browse files
README.md
CHANGED
|
@@ -125,13 +125,17 @@ The teacher is the engine's **own alpha-beta search** — the distillation
|
|
| 125 |
principle behind DeepMind's searchless grandmaster-level chess, at a size that
|
| 126 |
fits in L1 cache rather than a TPU pod. The student never searches.
|
| 127 |
|
| 128 |
-
- **Data**:
|
| 129 |
-
8 to 16 plies, labelled at 6k nodes/move by
|
| 130 |
-
|
| 131 |
-
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 135 |
- **Filtering**: positions are dropped when the side to move is in check or the
|
| 136 |
best move is a capture. There the tactic decides the game, not the static
|
| 137 |
evaluation, and training on them only teaches the network to imitate search —
|
|
|
|
| 125 |
principle behind DeepMind's searchless grandmaster-level chess, at a size that
|
| 126 |
fits in L1 cache rather than a TPU pod. The student never searches.
|
| 127 |
|
| 128 |
+
- **Data**: 10.2M positions from engine self-play out of randomised openings of
|
| 129 |
+
8 to 16 plies, labelled at 3k to 6k nodes/move, deduplicated by FEN across
|
| 130 |
+
every generation run ever made. The first two plies of real play are skipped —
|
| 131 |
+
those are the engine repairing whatever the random opening did.
|
| 132 |
+
|
| 133 |
+
This supersedes the previous release, which trained on 3.4M positions from a
|
| 134 |
+
single generation on the theory that one teacher beats an average of several.
|
| 135 |
+
Measured over 3000 games, that theory is worth **-15.5 Elo**: training on
|
| 136 |
+
everything, older labels included, beats training on the newest shard alone by
|
| 137 |
+
+15.5 with 95% confidence [+5.6, +25.5]. The older labels are weaker but they
|
| 138 |
+
are not noise, and there are seven million of them.
|
| 139 |
- **Filtering**: positions are dropped when the side to move is in check or the
|
| 140 |
best move is a capture. There the tactic decides the game, not the static
|
| 141 |
evaluation, and training on them only teaches the network to imitate search —
|