File size: 11,083 Bytes
4d66bc1
 
ddae13c
 
 
4d66bc1
 
 
 
 
 
 
 
 
 
 
 
ddae13c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4d66bc1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
---
license: mit
language:
  - en
  - fr
library_name: transformers
pipeline_tag: text-generation
tags:
  - chess
  - chess-engine
  - from-scratch
  - llama
  - small-model
---

# Philidor 51M

*[Version française plus bas](#philidor-51m-français)*

A language model trained **from scratch** to play chess, without ever being
given a single rule of the game.

It never sees a board. It does not know that pieces, squares or a king exist.
It receives a sequence of moves in UCI notation and predicts the next one,
exactly the way a language model predicts the next word.

> *"Pawns are the soul of chess."*
> François-André Danican Philidor, 1749

## Result

**97.86 % of the moves it proposes are legal**, in free generation, with no
constraint whatsoever. 95 % confidence interval: [97.65 – 98.05], measured on 20,000
positions from held-out games.

No rule was ever hard-coded. The model inferred the mechanics of the game from
790 million moves played by humans.

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("billygeekourson/philidor-51m")
model = AutoModelForCausalLM.from_pretrained("billygeekourson/philidor-51m")

ids = tok("e2e4 e7e5 g1f3 b8c6", return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=20, do_sample=True,
                     temperature=0.6, top_k=20, pad_token_id=0)
print(tok.decode(out[0], skip_special_tokens=True))
```

Moves are written in UCI, separated by spaces. **One move is exactly one
token**, so an 80-ply game takes 80 tokens.

## Playing without ever producing an illegal move

The bundled `vocab_uci.json` lets you mask impossible moves before choosing.
A single forward pass is enough, and the model keeps its preference ordering
among the playable moves.

```python
import json, torch, chess
from huggingface_hub import hf_hub_download

v = json.load(open(hf_hub_download("billygeekourson/philidor-51m",
                                   "vocab_uci.json")))
board = chess.Board()
board.push_uci("e2e4"); board.push_uci("c7c5")

logits = model(tok("e2e4 c7c5", return_tensors="pt").input_ids).logits[0, -1]
mask = torch.full_like(logits, float("-inf"))
for move in board.legal_moves:
    i = v["stoi"][move.uci()]
    mask[i] = logits[i]

print(v["itos"][int(mask.argmax())])   # g1f3
```

## Architecture

Decoder-only Transformer: 16 layers, width 512, 8 attention heads, hidden MLP 1408. Pre-norm, RMSNorm, RoPE, SwiGLU, tied
embeddings. Context of 256 moves, i.e. a whole game.

**51,397,120 non-embedding parameters** (52,406,272 total).

The architecture happens to match Llama exactly, which was not planned: the
four building blocks were each chosen on their own merits. The model therefore
loads as a standard `LlamaForCausalLM`, and the conversion was verified with a
maximum logit difference of `0.00e+00` at every sequence length.

## Vocabulary

**1,971 tokens**: the 1,968 geometrically possible UCI moves on a chessboard,
plus `<pad>`, `<bos>` and `<eos>`.

No BPE. The vocabulary is finite and known in advance, which makes legality
masking possible and spares the model from relearning that `e2` and `e4` form
a single unit.

## Data

One month of public Lichess archives, July 2026.

| | |
|---|---|
| Games read | 89,288,421 |
| Games kept | 11,035,777 (12.4 %) |
| Training tokens | 803,780,147 |

Filtering: both players between 1800 and 2600 Elo, bullet excluded, normal
termination, between 20 and 300 plies. The validation split is done **per
game**, never per token, so that no game is cut between the two sets.

## Training

A single RTX 3090. **2 hours**, 16,300 steps, one full epoch. bf16, AdamW,
cosine schedule with warmup, effective batch of 49,152 tokens. Measured MFU:
64 %.

Final loss: 1.6024 training, 1.6434 validation. Both curves stay
superimposed, so no overfitting.

## What the model knows

| Metric | Value |
|---|---|
| Legal moves in free generation | **97.86 %** |
| Agreement with the human move, top-1 | 51.42 % |
| Agreement with the human move, top-5 | 88.72 % |
| Castling | 100.00 % |
| En passant | 100.00 % |
| Promotion | 98.60 % |
| Getting out of check | 96.60 % |

All measurements are on held-out validation positions, at temperature 1.0 for
legality and on the argmax for the rest.

## Known limitations

**No board representation.** The model only understands a sequence of moves
from the initial position. It cannot resume from an arbitrary FEN position.

**About two moves in a hundred are illegal** without masking. For actual play, masking is essential.

**Modest playing strength.** On par with Stockfish capped at skill level 0, it
drops off at level 1. This is a search-free model: it plays the most probable
move after a single forward pass, with no lookahead.

**It plays endgames less well than openings.** Positions with few pieces offer
less statistical regularity to exploit.

## Related model

`philidor-142m`: same method, 142 M parameters, four months of data, twenty hours of training, 98.85 % legal moves. In a 400-game head-to-head it beats this model 305 to 32, a 290 Elo gap.

The interesting part: at equal data volume, tripling the model size barely changes legality, but gains two to three points of agreement with the human move. What drives playing strength is the amount of data, not model size.

## License

MIT for the model. Data comes from the public Lichess archives, released under
CC0.

---
---

# Philidor 51M (français)


Un modèle de langage entraîné **de zéro** à jouer aux échecs, sans qu'aucune
règle du jeu ne lui ait jamais été donnée.

Il ne voit pas d'échiquier. Il ne sait pas qu'il existe des pièces, des cases,
un roi. Il reçoit une suite de coups en notation UCI et prédit le suivant,
exactement comme un modèle de langage prédit le mot suivant.

> *« Les pions sont l'âme des échecs. »*
> François-André Danican Philidor, 1749

## Le résultat

**97,86 % des coups qu'il propose sont légaux**, en génération libre, sans
aucune contrainte. Intervalle de confiance à 95 % : [97,65 – 98,05], mesuré
sur 20 000 positions issues de parties de validation jamais vues.

Aucune règle n'a été codée. Le modèle a déduit la mécanique du jeu de
790 millions de coups joués par des humains.

## Utilisation

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("billygeekourson/philidor-51m")
model = AutoModelForCausalLM.from_pretrained("billygeekourson/philidor-51m")

ids = tok("e2e4 e7e5 g1f3 b8c6", return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=20, do_sample=True,
                     temperature=0.6, top_k=20, pad_token_id=0)
print(tok.decode(out[0], skip_special_tokens=True))
```

Les coups s'écrivent en UCI, séparés par des espaces. **Un coup vaut exactement
un token**, donc une partie de 80 demi-coups occupe 80 tokens.

## Jouer sans jamais produire de coup illégal

Le fichier `vocab_uci.json` permet de masquer les coups impossibles avant de
choisir. Un seul passage avant suffit, et le modèle conserve son ordre de
préférence entre les coups jouables.

```python
import json, torch, chess

v = json.load(open("vocab_uci.json"))
board = chess.Board()
board.push_uci("e2e4"); board.push_uci("c7c5")

logits = model(tok("e2e4 c7c5", return_tensors="pt").input_ids).logits[0, -1]
masque = torch.full_like(logits, float("-inf"))
for coup in board.legal_moves:
    i = v["stoi"][coup.uci()]
    masque[i] = logits[i]

print(v["itos"][int(masque.argmax())])   # g1f3
```

## Architecture

Transformer décodeur : 16 couches, dimension 512, 8 têtes d'attention, MLP
caché 1408. Pre-norm, RMSNorm, RoPE, SwiGLU, embeddings liés. Contexte de
256 coups, soit une partie entière.

**51 397 120 paramètres hors embeddings** (52 406 272 au total).

L'architecture correspond exactement à celle de Llama, ce qui n'était pas
prévu : les quatre briques ont été choisies séparément pour leurs mérites
propres. Le modèle se charge donc comme un `LlamaForCausalLM` standard, et la
conversion a été vérifiée : écart maximal de `0.00e+00` sur les logits, à
toutes les longueurs de séquence.

## Vocabulaire

**1 971 tokens** : les 1 968 coups UCI géométriquement possibles sur un
échiquier, plus `<pad>`, `<bos>` et `<eos>`.

Pas de BPE. Le vocabulaire est fini et connu d'avance, ce qui rend possible le
masquage de légalité et évite au modèle de réapprendre que `e2` et `e4`
forment une seule unité.

## Données

Un mois d'archives publiques Lichess, juillet 2026.

| | |
|---|---|
| Parties lues | 89 288 421 |
| Parties conservées | 11 035 777 (12,4 %) |
| Tokens d'entraînement | 803 780 147 |

Filtrage : les deux joueurs entre 1800 et 2600 Elo, bullet exclu, terminaison
normale, entre 20 et 300 demi-coups. Le découpage validation se fait par
partie et non par token, pour qu'aucune partie ne soit coupée entre les deux
jeux.

## Entraînement

Une seule RTX 3090. **2 heures**, 16 300 steps, une époque complète. bf16,
AdamW, cosinus avec chauffe, batch effectif de 49 152 tokens. MFU mesuré : 64 %.

Loss finale : 1,6024 en entraînement, 1,6434 en validation. Les deux courbes
restent superposées, donc aucun surapprentissage.

## Ce que le modèle sait

| Mesure | Valeur |
|---|---|
| Coups légaux en génération libre | **97,86 %** |
| Accord avec le coup humain, top-1 | 51,42 % |
| Accord avec le coup humain, top-5 | 88,72 % |
| Roque | 100,00 % |
| Prise en passant | 100,00 % |
| Promotion | 98,60 % |
| Sortie d'échec | 96,60 % |

Toutes les mesures portent sur des positions de validation jamais vues à
l'entraînement, à température 1,0 pour la légalité et sur l'argmax pour le
reste.

## Limites à connaître

**Aucune représentation de l'échiquier.** Le modèle ne comprend qu'une suite de
coups depuis la position initiale. Il ne sait pas repartir d'une position
arbitraire donnée en FEN.

**Environ deux coups sur cent sont illégaux** sans masquage. Pour jouer réellement,
le masquage est indispensable.

**Force de jeu modeste.** À égalité avec Stockfish bridé au niveau 0, il
décroche au niveau 1. C'est un modèle sans recherche : il joue le coup le plus
probable après un unique passage avant, sans aucune anticipation.

**Il ne joue pas les finales aussi bien que les ouvertures.** Les positions à
peu de pièces offrent moins de régularité statistique à exploiter.

## Modèle apparenté

`philidor-142m` : même méthode, 142 M de paramètres, quatre mois de données,
vingt heures d'entraînement, 98,85 % de coups légaux. En duel direct sur
400 parties il bat ce modèle-ci 305 à 32, soit un écart de 290 points d'Elo.

Le résultat est instructif : à volume de données égal, tripler la taille du
modèle ne change presque rien à la légalité, mais gagne deux à trois points
d'accord avec le coup humain. Ce qui fait la différence de force de jeu, c'est
le volume de données, pas la taille du modèle.

## Licence

MIT pour le modèle. Les données proviennent des archives publiques Lichess,
diffusées en CC0.