|
Download REPORT_S1.md from burrowdweller/mini-chessformer-v1: direct link, hf CLI and curl.
- Browser
- Download file 1.36 kB
-
https://huggingface.co/burrowdweller/mini-chessformer-v1/resolve/main/REPORT_S1.md
- Command line
-
hf download hf://burrowdweller/mini-chessformer-v1/REPORT_S1.md
-
curl -L -o REPORT_S1.md https://huggingface.co/burrowdweller/mini-chessformer-v1/resolve/main/REPORT_S1.md
1.36 kB
| # S1: vanilla PE vs GAB (matched steps) | |
| Date: 2026-08-03 | |
| Worktree: `.worktrees/chessformer-lite` | |
| Command: `python -m experiments.chessformer_lite.run_s1_compare --train-path /mnt/c/Users/jun/chessdb/data/train/data.parquet --out runs/chessformer_lite/s1 --steps 20000 --device cuda` | |
| Artifact: `runs/chessformer_lite/s1/compare.json` | |
| ## Setup | |
| | Knob | Value | | |
| |---|---| | |
| | steps | 20,000 | | |
| | batch | 256 | | |
| | warmup | 2,000 | | |
| | seed | 0 (shared trunk init matched vanilla→GAB) | | |
| | d_model / layers / heads / d_ff | 256 / 6 / 8 / 384 | | |
| | probe | 4,096 held-out train rows | | |
| ## Results | |
| | Arm | params | wall (s) | ms/step | final loss | policy CE (final) | probe top-1 | probe policy CE | train FLOPs | | |
| |---|---:|---:|---:|---:|---:|---:|---:|---:| | |
| | vanilla | 2,992,071 | 1074 | 53.7 | 3.134 | 2.455 | 0.464 | 1.758 | 5.97e15 | | |
| | **GAB** | 7,566,471 | 1171 | 58.6 | **2.776** | **2.127** | **0.488** | **1.626** | 1.51e16 | | |
| ms/step ratio GAB/vanilla = **1.09** (cap 1.30) → speed OK. | |
| ## Decision | |
| **GO — scale GAB.** | |
| - Better final loss and better probe policy CE. | |
| - Quality win under matched steps/data/seed with shared trunk init. | |
| - Step-time overhead well under 1.3× despite ~2.5× params (GAB templates dominate param count, not step cost). | |
| ## Next | |
| S2: `run_s2_scale --use-gab true` with flop_cap 4.665e17, then `eval_play` vs v1-final. | |