File size: 4,509 Bytes
c867bae
02c6c62
87f9a53
 
 
c867bae
 
87f9a53
c867bae
 
87f9a53
 
 
 
 
 
 
c867bae
 
02c6c62
87f9a53
 
 
724afb9
 
 
 
 
8c5a888
87f9a53
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
724afb9
8c5a888
 
724afb9
 
 
 
 
 
87f9a53
 
 
 
 
 
 
 
 
 
 
da1b5bb
87f9a53
724afb9
8c5a888
87f9a53
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
---
title: SLM Tetris Arena
emoji: 🧱
colorFrom: purple
colorTo: indigo
sdk: gradio
sdk_version: 6.28.0
python_version: "3.11"
app_file: app.py
pinned: false
license: apache-2.0
short_description: Tiny decoder-only LMs play Tetris zero-shot, ranked by Elo
tags:
  - leaderboard
  - evaluation
  - zero-shot
  - small-language-models
---

# 🧱 SLM Tetris Arena

Can a language model that only ever read text (e.g. FineWeb-edu) play Tetris **without any training on the game**?

Decoder-only models (50K–250M parameters, custom architectures welcome) play the **same piece sequence** side by
side on CPU. Two ways to play:

- **Match** (friendly): pick any 2+ models and the seed. Nothing is recorded.
- **Ranked**: press Play and the arena picks up to 4 models **at random** from its pool, all within 20M parameters of
  each other (models with 100M+ parameters can all play each other). The result updates a public **Elo** leaderboard.

## How a model plays

For each new piece the engine enumerates every legal placement (rotation × column, hard drop), simulates it and
describes the outcome in plain English. The model never sees the grid. It scores each description:

```
In Tetris, the goal is to clear lines, avoid holes and keep the stack low.
This move drops the piece into the lowest part of the board, clears one line, creates no new holes, keeps the stack low and leaves the surface flat.
It is a
```

value = `log P(" good move") − log P(" bad move")`. The highest value is played; exact ties are broken by a seeded
coin that is the same for every player. Because the value is a difference, a model's overall bias towards "good"
or "bad" cancels out, so no calibration is needed.

The better the model understands language (more/better pre-training), the better it should read the consequences
of each move. That's the hypothesis this arena tests.

- **Guided** protocol: the rules are stated in the prompt (reading comprehension).
- **Blind** protocol: `Here is a move from a game of Tetris.` Only pre-training knowledge.

Each protocol has its own leaderboard.

## Baselines

- 🎲 **Random**: uniform random placement (the floor).
- 📏 **Oracle reader**: ranks the same descriptions with fixed common sense (roughly the ceiling for a perfect reader).

## Elo

Only the Ranked tab changes Elo. The arena picks the players at random from the suggested models (models with fewer
ranked games are more likely to be picked), all within 20M parameters of each other, and uses a random seed (when
every picked model has 100M+ parameters, the 20M gap doesn't apply: big models can all play each other). Nobody
chooses who plays ranked, so Elo can't be farmed by pairing a model with weak opponents. The match runs on the server
in the background: it finishes and counts even if the viewer leaves, and only one ranked match runs at a time
(pressing Play while one is running lets you watch it).

Placement is decided by score (100/300/500/800 for 1–4 lines), then lines, then pieces survived (cap: 500 pieces).
Multiplayer Elo, K = 32: each pair of players is a game, scaled by 1/(N−1). Baselines are never rated.
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored in the
[results dataset](https://huggingface.co/datasets/DedeProGames/lm-tetris-arena-results).

## Configuration (Space variables / secrets)

| Name | Default | Purpose |
|---|---|---|
| `HF_TOKEN` (secret) | – | **Fine-grained** token with write access **only** to the results dataset. Without it, Elo lives in memory. |
| `RESULTS_REPO` | `DedeProGames/lm-tetris-arena-results` | Dataset that stores the leaderboard |
| `MAX_MODELS` | `4` | Language models per match |
| `MAX_PARAMS` | `250000000` | Parameter limit |
| `MIN_PARAMS` | `50000` | Smallest model allowed (smaller ones are rejected and removed from the leaderboard) |
| `MAX_PIECES` | `500` | Piece cap per game |
| `MAX_PARAM_GAP` | `20000000` | Largest size difference between models picked for a ranked match |
| `LARGE_FROM` | `100000000` | Models this size or bigger can play ranked against any other model this size (no gap) |
| `ALLOW_REMOTE_CODE` | `1` | Allow models with custom code (`trust_remote_code`) |
| `TORCH_THREADS` | `2` | CPU threads for inference |

⚠️ Custom-code models run arbitrary Python on this Space. The app removes `HF_TOKEN` from the environment before
loading any model, but use a fine-grained token scoped to the results dataset so the worst case is a revertible
commit on that dataset.