--- title: LM Tetris Arena emoji: 🧱 colorFrom: purple colorTo: indigo sdk: gradio sdk_version: 6.28.0 python_version: "3.11" app_file: app.py pinned: false license: apache-2.0 short_description: Tiny decoder-only LMs play Tetris zero-shot, ranked by Elo tags: - leaderboard - evaluation - zero-shot - small-language-models --- # 🧱 LM Tetris Arena Can a language model that only ever read text (e.g. FineWeb-edu) play Tetris **without any training on the game**? Pick 2+ decoder-only models (≤ 250M parameters, custom architectures welcome), watch them play the **same piece sequence** side by side on CPU, and see their **Elo** change on a public leaderboard. ## How a model plays For each new piece the engine enumerates every legal placement (rotation Ɨ column, hard drop), simulates it and describes the outcome in plain English. The model never sees the grid. It scores each description: ``` In Tetris, the goal is to clear lines, avoid holes and keep the stack low. This move drops the piece into the lowest part of the board, clears one line, creates no new holes, keeps the stack low and leaves the surface flat. It is a ``` value = `log P(" good move") āˆ’ log P(" bad move")`. The highest value is played; exact ties are broken by a seeded coin that is the same for every player. Because the value is a difference, a model's overall bias towards "good" or "bad" cancels out, so no calibration is needed. The better the model understands language (more/better pre-training), the better it should read the consequences of each move. That's the hypothesis this arena tests. - **Guided** protocol: the rules are stated in the prompt (reading comprehension). - **Blind** protocol: `Here is a move from a game of Tetris.` Only pre-training knowledge. Each protocol has its own leaderboard. ## Baselines - šŸŽ² **Random**: uniform random placement (the floor). - šŸ“ **Oracle reader**: ranks the same descriptions with fixed common sense (roughly the ceiling for a perfect reader). ## Elo Ranked matches use a random seed. Placement is decided by score (100/300/500/800 for 1–4 lines), then lines, then pieces survived (cap: 500 pieces). Multiplayer Elo, K = 32: each pair of players is a game, scaled by 1/(Nāˆ’1). Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored in the [results dataset](https://huggingface.co/datasets/DedeProGames/lm-tetris-arena-results). ## Configuration (Space variables / secrets) | Name | Default | Purpose | |---|---|---| | `HF_TOKEN` (secret) | – | **Fine-grained** token with write access **only** to the results dataset. Without it, Elo lives in memory. | | `RESULTS_REPO` | `DedeProGames/lm-tetris-arena-results` | Dataset that stores the leaderboard | | `MAX_MODELS` | `4` | Language models per match | | `MAX_PARAMS` | `250000000` | Parameter limit | | `MAX_PIECES` | `500` | Piece cap per game | | `ALLOW_REMOTE_CODE` | `1` | Allow models with custom code (`trust_remote_code`) | | `TORCH_THREADS` | `2` | CPU threads for inference | āš ļø Custom-code models run arbitrary Python on this Space. The app removes `HF_TOKEN` from the environment before loading any model, but use a fine-grained token scoped to the results dataset so the worst case is a revertible commit on that dataset.