--- language: en license: cc-by-nc-4.0 tags: - chess - chess-engine - transformer - from-scratch - small-model --- # Parallax-Chess-Preview A chess move prediction engine trained from scratch on a single laptop GPU. Predicts the best move given a board position using a custom chess tokenizer and a 4.4M parameter transformer. ## How It Works Unlike traditional chess engines (Stockfish, Leela), Parallax-Chess-Preview uses **pure neural network evaluation** — no hand-crafted rules, no opening books, no endgame tables. It learned chess purely from 50,000 Lichess puzzles. - **Custom chess tokenizer**: Board positions encoded as 66 tokens (64 squares + side-to-move + separator), moves encoded as (from-square, to-square) pairs. Total vocab: 402 tokens. - **Architecture**: Transformer decoder (6 layers, 256 dim, 4 heads, GQA 4:2) - **Training**: 100K steps on 50K Lichess puzzles, final loss 0.088 ## Benchmarks | Metric | Value | |--------|-------| | Legal move rate | 100% | | Speed | 6 moves/sec (CPU) | | Estimated ELO | ~700 | | vs Stockfish (max) | 0-1 (W), 1-0 (B, blunder) | | Training data | 50K Lichess puzzles | | Training time | ~1 hour on RTX 5060 | ## Quick Start ```python import chess from vela_chess_engine import VelaChessV2 engine = VelaChessV2("model.safetensors", "chess_encoding.py") board = chess.Board() move = engine.choose_move(board) print(move.uci()) # e.g. "g1f3" ``` ## Interactive Play ```bash python play.py ``` Then type moves in UCI format (e.g. `e2e4`). ## Architecture Details ``` SmallLMConfig( vocab_size=402, # Custom chess vocabulary d_model=256, # Embedding dimension n_heads=4, # Attention heads n_kv_heads=2, # GQA grouped query attention n_layers=6, # Transformer layers intermediate_size=1024, max_seq_len=128, norm_type='rms', rope_theta=10000.0, ) ``` ## Training Data 50,000 chess puzzles from the [Lichess puzzle database](https://database.lichess.org/#puzzles), filtered to ratings 800-2500 with 2-6 move solutions. ## Limitations - **~700 ELO** — plays at beginner level - No search tree — purely neural move prediction - Trained on puzzle data only (not full games) - Cannot play from arbitrary positions outside training distribution ## What This Proves Even a tiny 4.4M parameter model can: - Learn legal chess move patterns from data alone - Achieve 100% legal move rate - Make occasionally reasonable opening choices (Nf3, Nc3) - Be trained from scratch in 1 hour on a laptop GPU ## Hardware - **Training**: Single RTX 5060 Laptop GPU (8GB VRAM) - **Inference**: CPU-only, ~6 moves/sec ## License CC BY-NC 4.0 (weights), AGPL-3.0 (code)