billygeekourson commited on
Commit
ddae13c
·
verified ·
1 Parent(s): 4d66bc1

Fiche bilingue anglais / français

Browse files
Files changed (1) hide show
  1. README.md +155 -1
README.md CHANGED
@@ -1,6 +1,8 @@
1
  ---
2
  license: mit
3
- language: fr
 
 
4
  library_name: transformers
5
  pipeline_tag: text-generation
6
  tags:
@@ -13,6 +15,158 @@ tags:
13
 
14
  # Philidor 51M
15
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
16
  Un modèle de langage entraîné **de zéro** à jouer aux échecs, sans qu'aucune
17
  règle du jeu ne lui ait jamais été donnée.
18
 
 
1
  ---
2
  license: mit
3
+ language:
4
+ - en
5
+ - fr
6
  library_name: transformers
7
  pipeline_tag: text-generation
8
  tags:
 
15
 
16
  # Philidor 51M
17
 
18
+ *[Version française plus bas](#philidor-51m-français)*
19
+
20
+ A language model trained **from scratch** to play chess, without ever being
21
+ given a single rule of the game.
22
+
23
+ It never sees a board. It does not know that pieces, squares or a king exist.
24
+ It receives a sequence of moves in UCI notation and predicts the next one,
25
+ exactly the way a language model predicts the next word.
26
+
27
+ > *"Pawns are the soul of chess."*
28
+ > François-André Danican Philidor, 1749
29
+
30
+ ## Result
31
+
32
+ **97.86 % of the moves it proposes are legal**, in free generation, with no
33
+ constraint whatsoever. 95 % confidence interval: [97.65 – 98.05], measured on 20,000
34
+ positions from held-out games.
35
+
36
+ No rule was ever hard-coded. The model inferred the mechanics of the game from
37
+ 790 million moves played by humans.
38
+
39
+ ## Usage
40
+
41
+ ```python
42
+ from transformers import AutoModelForCausalLM, AutoTokenizer
43
+
44
+ tok = AutoTokenizer.from_pretrained("billygeekourson/philidor-51m")
45
+ model = AutoModelForCausalLM.from_pretrained("billygeekourson/philidor-51m")
46
+
47
+ ids = tok("e2e4 e7e5 g1f3 b8c6", return_tensors="pt").input_ids
48
+ out = model.generate(ids, max_new_tokens=20, do_sample=True,
49
+ temperature=0.6, top_k=20, pad_token_id=0)
50
+ print(tok.decode(out[0], skip_special_tokens=True))
51
+ ```
52
+
53
+ Moves are written in UCI, separated by spaces. **One move is exactly one
54
+ token**, so an 80-ply game takes 80 tokens.
55
+
56
+ ## Playing without ever producing an illegal move
57
+
58
+ The bundled `vocab_uci.json` lets you mask impossible moves before choosing.
59
+ A single forward pass is enough, and the model keeps its preference ordering
60
+ among the playable moves.
61
+
62
+ ```python
63
+ import json, torch, chess
64
+ from huggingface_hub import hf_hub_download
65
+
66
+ v = json.load(open(hf_hub_download("billygeekourson/philidor-51m",
67
+ "vocab_uci.json")))
68
+ board = chess.Board()
69
+ board.push_uci("e2e4"); board.push_uci("c7c5")
70
+
71
+ logits = model(tok("e2e4 c7c5", return_tensors="pt").input_ids).logits[0, -1]
72
+ mask = torch.full_like(logits, float("-inf"))
73
+ for move in board.legal_moves:
74
+ i = v["stoi"][move.uci()]
75
+ mask[i] = logits[i]
76
+
77
+ print(v["itos"][int(mask.argmax())]) # g1f3
78
+ ```
79
+
80
+ ## Architecture
81
+
82
+ Decoder-only Transformer: 16 layers, width 512, 8 attention heads, hidden MLP 1408. Pre-norm, RMSNorm, RoPE, SwiGLU, tied
83
+ embeddings. Context of 256 moves, i.e. a whole game.
84
+
85
+ **51,397,120 non-embedding parameters** (52,406,272 total).
86
+
87
+ The architecture happens to match Llama exactly, which was not planned: the
88
+ four building blocks were each chosen on their own merits. The model therefore
89
+ loads as a standard `LlamaForCausalLM`, and the conversion was verified with a
90
+ maximum logit difference of `0.00e+00` at every sequence length.
91
+
92
+ ## Vocabulary
93
+
94
+ **1,971 tokens**: the 1,968 geometrically possible UCI moves on a chessboard,
95
+ plus `<pad>`, `<bos>` and `<eos>`.
96
+
97
+ No BPE. The vocabulary is finite and known in advance, which makes legality
98
+ masking possible and spares the model from relearning that `e2` and `e4` form
99
+ a single unit.
100
+
101
+ ## Data
102
+
103
+ One month of public Lichess archives, July 2026.
104
+
105
+ | | |
106
+ |---|---|
107
+ | Games read | 89,288,421 |
108
+ | Games kept | 11,035,777 (12.4 %) |
109
+ | Training tokens | 803,780,147 |
110
+
111
+ Filtering: both players between 1800 and 2600 Elo, bullet excluded, normal
112
+ termination, between 20 and 300 plies. The validation split is done **per
113
+ game**, never per token, so that no game is cut between the two sets.
114
+
115
+ ## Training
116
+
117
+ A single RTX 3090. **2 hours**, 16,300 steps, one full epoch. bf16, AdamW,
118
+ cosine schedule with warmup, effective batch of 49,152 tokens. Measured MFU:
119
+ 64 %.
120
+
121
+ Final loss: 1.6024 training, 1.6434 validation. Both curves stay
122
+ superimposed, so no overfitting.
123
+
124
+ ## What the model knows
125
+
126
+ | Metric | Value |
127
+ |---|---|
128
+ | Legal moves in free generation | **97.86 %** |
129
+ | Agreement with the human move, top-1 | 51.42 % |
130
+ | Agreement with the human move, top-5 | 88.72 % |
131
+ | Castling | 100.00 % |
132
+ | En passant | 100.00 % |
133
+ | Promotion | 98.60 % |
134
+ | Getting out of check | 96.60 % |
135
+
136
+ All measurements are on held-out validation positions, at temperature 1.0 for
137
+ legality and on the argmax for the rest.
138
+
139
+ ## Known limitations
140
+
141
+ **No board representation.** The model only understands a sequence of moves
142
+ from the initial position. It cannot resume from an arbitrary FEN position.
143
+
144
+ **About two moves in a hundred are illegal** without masking. For actual play, masking is essential.
145
+
146
+ **Modest playing strength.** On par with Stockfish capped at skill level 0, it
147
+ drops off at level 1. This is a search-free model: it plays the most probable
148
+ move after a single forward pass, with no lookahead.
149
+
150
+ **It plays endgames less well than openings.** Positions with few pieces offer
151
+ less statistical regularity to exploit.
152
+
153
+ ## Related model
154
+
155
+ `philidor-142m`: same method, 142 M parameters, four months of data, twenty hours of training, 98.85 % legal moves. In a 400-game head-to-head it beats this model 305 to 32, a 290 Elo gap.
156
+
157
+ The interesting part: at equal data volume, tripling the model size barely changes legality, but gains two to three points of agreement with the human move. What drives playing strength is the amount of data, not model size.
158
+
159
+ ## License
160
+
161
+ MIT for the model. Data comes from the public Lichess archives, released under
162
+ CC0.
163
+
164
+ ---
165
+ ---
166
+
167
+ # Philidor 51M (français)
168
+
169
+
170
  Un modèle de langage entraîné **de zéro** à jouer aux échecs, sans qu'aucune
171
  règle du jeu ne lui ait jamais été donnée.
172