mini-chessformer-v1 / REPORT_S1.md
burrowdweller's picture
Upload mini-chessformer-v1 (ONNX + PT + standalone inference)
837e238 verified
|
Raw History Blame Contribute Delete
1.36 kB

S1: vanilla PE vs GAB (matched steps)

Date: 2026-08-03
Worktree: .worktrees/chessformer-lite
Command: python -m experiments.chessformer_lite.run_s1_compare --train-path /mnt/c/Users/jun/chessdb/data/train/data.parquet --out runs/chessformer_lite/s1 --steps 20000 --device cuda
Artifact: runs/chessformer_lite/s1/compare.json

Setup

Knob Value
steps 20,000
batch 256
warmup 2,000
seed 0 (shared trunk init matched vanilla→GAB)
d_model / layers / heads / d_ff 256 / 6 / 8 / 384
probe 4,096 held-out train rows

Results

Arm params wall (s) ms/step final loss policy CE (final) probe top-1 probe policy CE train FLOPs
vanilla 2,992,071 1074 53.7 3.134 2.455 0.464 1.758 5.97e15
GAB 7,566,471 1171 58.6 2.776 2.127 0.488 1.626 1.51e16

ms/step ratio GAB/vanilla = 1.09 (cap 1.30) → speed OK.

Decision

GO — scale GAB.

  • Better final loss and better probe policy CE.
  • Quality win under matched steps/data/seed with shared trunk init.
  • Step-time overhead well under 1.3× despite ~2.5× params (GAB templates dominate param count, not step cost).

Next

S2: run_s2_scale --use-gab true with flop_cap 4.665e17, then eval_play vs v1-final.