ParallaxOpen's picture
Upload README.md with huggingface_hub
92fccf2 verified
|
Raw History Blame
2.7 kB
---
language: en
license: cc-by-nc-4.0
tags:
- chess
- chess-engine
- transformer
- from-scratch
- small-model
---
# Parallax-Chess-Preview
A chess move prediction engine trained from scratch on a single laptop GPU. Predicts the best move given a board position using a custom chess tokenizer and a 4.4M parameter transformer.
## How It Works
Unlike traditional chess engines (Stockfish, Leela), Parallax-Chess-Preview uses **pure neural network evaluation** — no hand-crafted rules, no opening books, no endgame tables. It learned chess purely from 50,000 Lichess puzzles.
- **Custom chess tokenizer**: Board positions encoded as 66 tokens (64 squares + side-to-move + separator), moves encoded as (from-square, to-square) pairs. Total vocab: 402 tokens.
- **Architecture**: Transformer decoder (6 layers, 256 dim, 4 heads, GQA 4:2)
- **Training**: 100K steps on 50K Lichess puzzles, final loss 0.088
## Benchmarks
| Metric | Value |
|--------|-------|
| Legal move rate | 100% |
| Speed | 6 moves/sec (CPU) |
| Estimated ELO | ~700 |
| vs Stockfish (max) | 0-1 (W), 1-0 (B, blunder) |
| Training data | 50K Lichess puzzles |
| Training time | ~1 hour on RTX 5060 |
## Quick Start
```python
import chess
from vela_chess_engine import VelaChessV2
engine = VelaChessV2("model.safetensors", "chess_encoding.py")
board = chess.Board()
move = engine.choose_move(board)
print(move.uci()) # e.g. "g1f3"
```
## Interactive Play
```bash
python play.py
```
Then type moves in UCI format (e.g. `e2e4`).
## Architecture Details
```
SmallLMConfig(
vocab_size=402, # Custom chess vocabulary
d_model=256, # Embedding dimension
n_heads=4, # Attention heads
n_kv_heads=2, # GQA grouped query attention
n_layers=6, # Transformer layers
intermediate_size=1024,
max_seq_len=128,
norm_type='rms',
rope_theta=10000.0,
)
```
## Training Data
50,000 chess puzzles from the [Lichess puzzle database](https://database.lichess.org/#puzzles), filtered to ratings 800-2500 with 2-6 move solutions.
## Limitations
- **~700 ELO** — plays at beginner level
- No search tree — purely neural move prediction
- Trained on puzzle data only (not full games)
- Cannot play from arbitrary positions outside training distribution
## What This Proves
Even a tiny 4.4M parameter model can:
- Learn legal chess move patterns from data alone
- Achieve 100% legal move rate
- Make occasionally reasonable opening choices (Nf3, Nc3)
- Be trained from scratch in 1 hour on a laptop GPU
## Hardware
- **Training**: Single RTX 5060 Laptop GPU (8GB VRAM)
- **Inference**: CPU-only, ~6 moves/sec
## License
CC BY-NC 4.0 (weights), AGPL-3.0 (code)