shubhxho commited on
Commit
20b258b
·
verified ·
1 Parent(s): c19d0c6

Document the +15.5 Elo network and its data

Browse files
Files changed (1) hide show
  1. README.md +11 -7
README.md CHANGED
@@ -125,13 +125,17 @@ The teacher is the engine's **own alpha-beta search** — the distillation
125
  principle behind DeepMind's searchless grandmaster-level chess, at a size that
126
  fits in L1 cache rather than a TPU pod. The student never searches.
127
 
128
- - **Data**: 3.4M positions from engine self-play out of randomised openings of
129
- 8 to 16 plies, labelled at 6k nodes/move by the current engine. Positions are
130
- deduplicated by Zobrist key across the whole run, and the first two plies of
131
- real play are skipped — those are the engine repairing whatever the random
132
- opening did. Earlier releases used 3.36M positions labelled at 3k and 5k
133
- nodes by weaker versions of the same search; those shards are kept, not
134
- mixed in, so each network has one teacher rather than an average of several.
 
 
 
 
135
  - **Filtering**: positions are dropped when the side to move is in check or the
136
  best move is a capture. There the tactic decides the game, not the static
137
  evaluation, and training on them only teaches the network to imitate search —
 
125
  principle behind DeepMind's searchless grandmaster-level chess, at a size that
126
  fits in L1 cache rather than a TPU pod. The student never searches.
127
 
128
+ - **Data**: 10.2M positions from engine self-play out of randomised openings of
129
+ 8 to 16 plies, labelled at 3k to 6k nodes/move, deduplicated by FEN across
130
+ every generation run ever made. The first two plies of real play are skipped —
131
+ those are the engine repairing whatever the random opening did.
132
+
133
+ This supersedes the previous release, which trained on 3.4M positions from a
134
+ single generation on the theory that one teacher beats an average of several.
135
+ Measured over 3000 games, that theory is worth **-15.5 Elo**: training on
136
+ everything, older labels included, beats training on the newest shard alone by
137
+ +15.5 with 95% confidence [+5.6, +25.5]. The older labels are weaker but they
138
+ are not noise, and there are seven million of them.
139
  - **Filtering**: positions are dropped when the side to move is in check or the
140
  best move is a capture. There the tactic decides the game, not the static
141
  evaluation, and training on them only teaches the network to imitate search —