DedeProGames commited on
Commit
724afb9
·
verified ·
1 Parent(s): a8c67d3

Document Match and Ranked tabs

Browse files
Files changed (1) hide show
  1. README.md +15 -8
README.md CHANGED
@@ -21,8 +21,12 @@ tags:
21
 
22
  Can a language model that only ever read text (e.g. FineWeb-edu) play Tetris **without any training on the game**?
23
 
24
- Pick 2+ decoder-only models (50K–250M parameters, custom architectures welcome), watch them play the **same piece
25
- sequence** side by side on CPU, and see their **Elo** change on a public leaderboard.
 
 
 
 
26
 
27
  ## How a model plays
28
 
@@ -54,11 +58,14 @@ Each protocol has its own leaderboard.
54
 
55
  ## Elo
56
 
57
- Ranked matches use a random seed. Placement is decided by score (100/300/500/800 for 1–4 lines), then lines, then
58
- pieces survived (cap: 500 pieces). Multiplayer Elo, K = 32: each pair of players is a game, scaled by 1/(N−1).
59
- Only language models whose sizes are within 20M parameters of each other can play ranked (so a big model can't
60
- farm Elo from tiny ones); a ranked match outside that range is blocked (it can still be played unranked).
61
- Baselines are never rated.
 
 
 
62
  Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored in the
63
  [results dataset](https://huggingface.co/datasets/DedeProGames/lm-tetris-arena-results).
64
 
@@ -72,7 +79,7 @@ Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored
72
  | `MAX_PARAMS` | `250000000` | Parameter limit |
73
  | `MIN_PARAMS` | `50000` | Smallest model allowed (smaller ones are rejected and removed from the leaderboard) |
74
  | `MAX_PIECES` | `500` | Piece cap per game |
75
- | `MAX_PARAM_GAP` | `20000000` | Largest size difference allowed in a ranked match |
76
  | `ALLOW_REMOTE_CODE` | `1` | Allow models with custom code (`trust_remote_code`) |
77
  | `TORCH_THREADS` | `2` | CPU threads for inference |
78
 
 
21
 
22
  Can a language model that only ever read text (e.g. FineWeb-edu) play Tetris **without any training on the game**?
23
 
24
+ Decoder-only models (50K–250M parameters, custom architectures welcome) play the **same piece sequence** side by
25
+ side on CPU. Two ways to play:
26
+
27
+ - **Match** (friendly): pick any 2+ models and the seed. Nothing is recorded.
28
+ - **Ranked**: press Play and the arena picks up to 4 models **at random** from its pool, all within 20M parameters of
29
+ each other. The result updates a public **Elo** leaderboard.
30
 
31
  ## How a model plays
32
 
 
58
 
59
  ## Elo
60
 
61
+ Only the Ranked tab changes Elo. The arena picks the players at random from the suggested models (models with fewer
62
+ ranked games are more likely to be picked), all within 20M parameters of each other, and uses a random seed. Nobody
63
+ chooses who plays ranked, so Elo can't be farmed by pairing a model with weak opponents. The match runs on the server
64
+ in the background: it finishes and counts even if the viewer leaves, and only one ranked match runs at a time
65
+ (pressing Play while one is running lets you watch it).
66
+
67
+ Placement is decided by score (100/300/500/800 for 1–4 lines), then lines, then pieces survived (cap: 500 pieces).
68
+ Multiplayer Elo, K = 32: each pair of players is a game, scaled by 1/(N−1). Baselines are never rated.
69
  Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored in the
70
  [results dataset](https://huggingface.co/datasets/DedeProGames/lm-tetris-arena-results).
71
 
 
79
  | `MAX_PARAMS` | `250000000` | Parameter limit |
80
  | `MIN_PARAMS` | `50000` | Smallest model allowed (smaller ones are rejected and removed from the leaderboard) |
81
  | `MAX_PIECES` | `500` | Piece cap per game |
82
+ | `MAX_PARAM_GAP` | `20000000` | Largest size difference between models picked for a ranked match |
83
  | `ALLOW_REMOTE_CODE` | `1` | Allow models with custom code (`trust_remote_code`) |
84
  | `TORCH_THREADS` | `2` | CPU threads for inference |
85