File size: 5,809 Bytes
613f492 31971bb 613f492 31971bb 613f492 31971bb 613f492 31971bb 613f492 de5529e 31971bb 613f492 31971bb 613f492 31971bb 613f492 31971bb 613f492 31971bb 613f492 31971bb 613f492 31971bb 5c6b472 31971bb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 | ---
license: mit
tags:
- onepass
- connect-four
- tiny-model
- on-device
- onnx
datasets:
- TonyCWang/ConnectFour
---
# onepass-c4 β a one-pass specialist that plays Connect Four
A *one-pass specialist* scores a supplied list of options in one forward pass: context bytes in, one
score per option out, no generation, no tokenizer. This one plays Connect Four: the board goes in as
44 bytes, the legal columns go in as short strings (`column 4`), and the highest-scoring column is
played. **No search, no hand-written rules** β the scores are the whole decision.
**Play it in your browser:** https://precisit.github.io/onepass-web/demo/c4/ (the page's arena pits v2
against v1 or a random player).
**The story behind the model:** [Meet your one-pass AI opponent](https://precisit.com/en/blog/onepass-connect-four/)
follows the first attempt, the better teacher and the browser release.
## v2 (current)
| | |
| --- | --- |
| architecture | the one-pass option scorer from [one-pass-specialists](https://github.com/precisit/one-pass-specialists) (byte embedding β 8 transformer layers, width 256 β option attention head) β the same architecture as v1, larger config |
| parameters | 7.38 M |
| input | context 44 bytes: `1:` or `2:` (did the player to move open the game?) + 42 cells, bottom row first, `m` = mover, `t` = opponent, `.` = empty; options `column 1` β¦ `column 7`, 8 bytes each |
| files | `onepass-c4-v2-int8.onnx` (7.8 MB, ONNX Runtime dynamic int8 β what the demo loads) Β· `onepass-c4-v2.onnx` (fp32, 29.7 MB) Β· `v2/model.safetensors` + `v2/config.json` (toolkit checkpoint) |
| ONNX inputs | `context_ids` int32 [1, 44], `option_ids` int32 [1, 7, 8], `option_mask` int32 [1, 7]; bytes are UTF-8 + 1, 0 = padding; output `logits` [1, 7] |
### How well it plays
Frozen protocol, pre-registered before training: 200 games per match, colours alternate, games start
from the empty board, and **both** players make a uniformly random move 5 % of the time (otherwise two
deterministic players replay one game). Score = (wins + draws/2) / games, 95 % Wilson intervals.
| opponent | v2 score (checkpoint) | v2 score (int8 file) | v1 score |
| --- | ---: | ---: | ---: |
| depth-2 alpha-beta bot | 0.905 [0.856β0.938] | 0.910 | 0.03 |
| depth-4 alpha-beta bot | **0.915** [0.868β0.946] | 0.905 | 0.03 |
| depth-6 alpha-beta bot | **0.893** [0.842β0.928] | 0.878 | 0.02 |
| perfect player (exact solver) | 0.475 [0.407β0.544] | β | β |
| random player | 1.000 | β | 0.885 |
| v1 | 0.985 | β | β |
For scale, a *perfect* player scores 0.890 against the depth-4 bot and 0.925 against the depth-6 bot
under the same protocol β its own 5 % random moves cost it the rest. Against the perfect player
itself v2 is close to even.
Move level, on a held-out set of 17 325 positions covering every ply (half from strong self-play, half
from games between weak and strong players): **98.7 %** of v2's moves keep the game-theoretic value
of the position (win stays win, draw stays draw); 97.0 % on positions where the choice matters; 93.5 %
of its moves are a fastest win / slowest loss. It reads the board: reassigning every stone's owner
(column heights kept) drops the non-trivial rate by 49 points.
Browser: the int8 file runs in ONNX Runtime Web; ~20 ms per decision on wasm and ~9 ms on WebNN in
headless Chromium on an M1 Max (v1 on the same machine: 10.8 ms on wasm).
### How it was trained
Full recipe, protocol and every measurement: [`examples/c4`](https://github.com/precisit/one-pass-specialists/tree/main/examples/c4)
in one-pass-specialists.
* **Labels:** every legal column of every position scored exactly (win/draw/loss and distance), by
[connect-four-ai](https://github.com/benjaminrall/connect-four-ai) (MIT), verified before use
against Pascal Pons' public test sets (6 000 positions), the UCI Connect-4 8-ply database
(67 557 positions) and an independent solver β zero disagreements.
* **Positions:** 41.6 M unique positions from [TonyCWang/ConnectFour](https://huggingface.co/datasets/TonyCWang/ConnectFour)
(MIT; solver self-play at varied temperature, every column scored β re-solved on a 20 k sample with
zero disagreements) plus 513 k positions from our own games between weak and strong players.
* **Target:** listwise cross-entropy over the legal columns, towards a softmax of the exact scores in
which the win/draw/loss class dominates and distance breaks ties.
* **Schedule:** 18 000 steps Γ 1 024 positions (β 18 M samples), ply-balanced sampling, random
horizontal mirroring; about two hours on one Apple M5 Pro.
### Limitations
* Not perfect: 1.3 % of its moves give away the game-theoretic value, and it loses to a perfect player
about as often as it wins under the noisy protocol.
* Its choice between near-equal columns is not mirror-symmetric (~77 % consistency on mirrored boards).
* The int8 file picks a different column than the fp32 one on 1.9 % of positions (mostly near-ties; its value-preserving rate is the same, 98.75 %); its measured strength is
shown separately above.
## v1 (previous, kept for comparison)
`onepass-c4-8x24.onnx` + `model.safetensors` (706 k parameters, 224 context bytes, 8 Γ 24 option
bytes): trained on ~80 k endgame positions only. It read the position from a move-history line
rather than the board, and scores 0.02β0.03 against the search bots above. It stays here because the
demo's arena plays it against v2.
## License and credits
MIT. Labels by connect-four-ai (MIT, Benjamin Rall); positions from TonyCWang/ConnectFour (MIT);
evaluation positions from the UCI Connect-4 database (John Tromp, CC BY 4.0) and Pascal Pons' test
sets were used only to verify the labeller. Architecture adapted in one-pass-specialists from
Cua-S1 / jevlike (MIT).
|