ParallaxOpen's picture
Upload README.md with huggingface_hub
92fccf2 verified
|
Raw History Blame
2.7 kB
metadata
language: en
license: cc-by-nc-4.0
tags:
  - chess
  - chess-engine
  - transformer
  - from-scratch
  - small-model

Parallax-Chess-Preview

A chess move prediction engine trained from scratch on a single laptop GPU. Predicts the best move given a board position using a custom chess tokenizer and a 4.4M parameter transformer.

How It Works

Unlike traditional chess engines (Stockfish, Leela), Parallax-Chess-Preview uses pure neural network evaluation — no hand-crafted rules, no opening books, no endgame tables. It learned chess purely from 50,000 Lichess puzzles.

  • Custom chess tokenizer: Board positions encoded as 66 tokens (64 squares + side-to-move + separator), moves encoded as (from-square, to-square) pairs. Total vocab: 402 tokens.
  • Architecture: Transformer decoder (6 layers, 256 dim, 4 heads, GQA 4:2)
  • Training: 100K steps on 50K Lichess puzzles, final loss 0.088

Benchmarks

Metric Value
Legal move rate 100%
Speed 6 moves/sec (CPU)
Estimated ELO ~700
vs Stockfish (max) 0-1 (W), 1-0 (B, blunder)
Training data 50K Lichess puzzles
Training time ~1 hour on RTX 5060

Quick Start

import chess
from vela_chess_engine import VelaChessV2

engine = VelaChessV2("model.safetensors", "chess_encoding.py")
board = chess.Board()
move = engine.choose_move(board)
print(move.uci())  # e.g. "g1f3"

Interactive Play

python play.py

Then type moves in UCI format (e.g. e2e4).

Architecture Details

SmallLMConfig(
  vocab_size=402,       # Custom chess vocabulary
  d_model=256,          # Embedding dimension
  n_heads=4,            # Attention heads
  n_kv_heads=2,         # GQA grouped query attention
  n_layers=6,           # Transformer layers
  intermediate_size=1024,
  max_seq_len=128,
  norm_type='rms',
  rope_theta=10000.0,
)

Training Data

50,000 chess puzzles from the Lichess puzzle database, filtered to ratings 800-2500 with 2-6 move solutions.

Limitations

  • ~700 ELO — plays at beginner level
  • No search tree — purely neural move prediction
  • Trained on puzzle data only (not full games)
  • Cannot play from arbitrary positions outside training distribution

What This Proves

Even a tiny 4.4M parameter model can:

  • Learn legal chess move patterns from data alone
  • Achieve 100% legal move rate
  • Make occasionally reasonable opening choices (Nf3, Nc3)
  • Be trained from scratch in 1 hour on a laptop GPU

Hardware

  • Training: Single RTX 5060 Laptop GPU (8GB VRAM)
  • Inference: CPU-only, ~6 moves/sec

License

CC BY-NC 4.0 (weights), AGPL-3.0 (code)