Spaces:
Running
Running
Document Match and Ranked tabs
Browse files
README.md
CHANGED
|
@@ -21,8 +21,12 @@ tags:
|
|
| 21 |
|
| 22 |
Can a language model that only ever read text (e.g. FineWeb-edu) play Tetris **without any training on the game**?
|
| 23 |
|
| 24 |
-
|
| 25 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 26 |
|
| 27 |
## How a model plays
|
| 28 |
|
|
@@ -54,11 +58,14 @@ Each protocol has its own leaderboard.
|
|
| 54 |
|
| 55 |
## Elo
|
| 56 |
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
|
|
|
|
|
|
|
|
|
| 62 |
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored in the
|
| 63 |
[results dataset](https://huggingface.co/datasets/DedeProGames/lm-tetris-arena-results).
|
| 64 |
|
|
@@ -72,7 +79,7 @@ Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored
|
|
| 72 |
| `MAX_PARAMS` | `250000000` | Parameter limit |
|
| 73 |
| `MIN_PARAMS` | `50000` | Smallest model allowed (smaller ones are rejected and removed from the leaderboard) |
|
| 74 |
| `MAX_PIECES` | `500` | Piece cap per game |
|
| 75 |
-
| `MAX_PARAM_GAP` | `20000000` | Largest size difference
|
| 76 |
| `ALLOW_REMOTE_CODE` | `1` | Allow models with custom code (`trust_remote_code`) |
|
| 77 |
| `TORCH_THREADS` | `2` | CPU threads for inference |
|
| 78 |
|
|
|
|
| 21 |
|
| 22 |
Can a language model that only ever read text (e.g. FineWeb-edu) play Tetris **without any training on the game**?
|
| 23 |
|
| 24 |
+
Decoder-only models (50K–250M parameters, custom architectures welcome) play the **same piece sequence** side by
|
| 25 |
+
side on CPU. Two ways to play:
|
| 26 |
+
|
| 27 |
+
- **Match** (friendly): pick any 2+ models and the seed. Nothing is recorded.
|
| 28 |
+
- **Ranked**: press Play and the arena picks up to 4 models **at random** from its pool, all within 20M parameters of
|
| 29 |
+
each other. The result updates a public **Elo** leaderboard.
|
| 30 |
|
| 31 |
## How a model plays
|
| 32 |
|
|
|
|
| 58 |
|
| 59 |
## Elo
|
| 60 |
|
| 61 |
+
Only the Ranked tab changes Elo. The arena picks the players at random from the suggested models (models with fewer
|
| 62 |
+
ranked games are more likely to be picked), all within 20M parameters of each other, and uses a random seed. Nobody
|
| 63 |
+
chooses who plays ranked, so Elo can't be farmed by pairing a model with weak opponents. The match runs on the server
|
| 64 |
+
in the background: it finishes and counts even if the viewer leaves, and only one ranked match runs at a time
|
| 65 |
+
(pressing Play while one is running lets you watch it).
|
| 66 |
+
|
| 67 |
+
Placement is decided by score (100/300/500/800 for 1–4 lines), then lines, then pieces survived (cap: 500 pieces).
|
| 68 |
+
Multiplayer Elo, K = 32: each pair of players is a game, scaled by 1/(N−1). Baselines are never rated.
|
| 69 |
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored in the
|
| 70 |
[results dataset](https://huggingface.co/datasets/DedeProGames/lm-tetris-arena-results).
|
| 71 |
|
|
|
|
| 79 |
| `MAX_PARAMS` | `250000000` | Parameter limit |
|
| 80 |
| `MIN_PARAMS` | `50000` | Smallest model allowed (smaller ones are rejected and removed from the leaderboard) |
|
| 81 |
| `MAX_PIECES` | `500` | Piece cap per game |
|
| 82 |
+
| `MAX_PARAM_GAP` | `20000000` | Largest size difference between models picked for a ranked match |
|
| 83 |
| `ALLOW_REMOTE_CODE` | `1` | Allow models with custom code (`trust_remote_code`) |
|
| 84 |
| `TORCH_THREADS` | `2` | CPU threads for inference |
|
| 85 |
|