Spaces:
Running
Running
Document ranked without size limit
Browse files
README.md
CHANGED
|
@@ -25,8 +25,8 @@ Decoder-only models (50K–250M parameters, custom architectures welcome) play t
|
|
| 25 |
side on CPU. Two ways to play:
|
| 26 |
|
| 27 |
- **Match** (friendly): pick any 2+ models and the seed. Nothing is recorded.
|
| 28 |
-
- **Ranked**: press Play and the arena picks up to 4 models **at random** from its pool,
|
| 29 |
-
|
| 30 |
|
| 31 |
## How a model plays
|
| 32 |
|
|
@@ -59,9 +59,10 @@ Each protocol has its own leaderboard.
|
|
| 59 |
## Elo
|
| 60 |
|
| 61 |
Only the Ranked tab changes Elo. The arena picks the players at random from the suggested models (models with fewer
|
| 62 |
-
ranked games are more likely to be picked),
|
| 63 |
-
|
| 64 |
-
|
|
|
|
| 65 |
in the background: it finishes and counts even if the viewer leaves, and only one ranked match runs at a time
|
| 66 |
(pressing Play while one is running lets you watch it).
|
| 67 |
|
|
@@ -80,8 +81,8 @@ Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored
|
|
| 80 |
| `MAX_PARAMS` | `250000000` | Parameter limit |
|
| 81 |
| `MIN_PARAMS` | `50000` | Smallest model allowed (smaller ones are rejected and removed from the leaderboard) |
|
| 82 |
| `MAX_PIECES` | `500` | Piece cap per game |
|
| 83 |
-
| `MAX_PARAM_GAP` | `
|
| 84 |
-
| `LARGE_FROM` | `100000000` |
|
| 85 |
| `ALLOW_REMOTE_CODE` | `1` | Allow models with custom code (`trust_remote_code`) |
|
| 86 |
| `TORCH_THREADS` | `2` | CPU threads for inference |
|
| 87 |
|
|
|
|
| 25 |
side on CPU. Two ways to play:
|
| 26 |
|
| 27 |
- **Match** (friendly): pick any 2+ models and the seed. Nothing is recorded.
|
| 28 |
+
- **Ranked**: press Play and the arena picks up to 4 models **at random** from its pool, of any size. The result updates
|
| 29 |
+
a public **Elo** leaderboard.
|
| 30 |
|
| 31 |
## How a model plays
|
| 32 |
|
|
|
|
| 59 |
## Elo
|
| 60 |
|
| 61 |
Only the Ranked tab changes Elo. The arena picks the players at random from the suggested models (models with fewer
|
| 62 |
+
ranked games are more likely to be picked), of any size, and uses a random seed. Nobody chooses who plays ranked, so
|
| 63 |
+
Elo can't be farmed by pairing a model with weak opponents. Any model can meet any other, so every rating sits on one
|
| 64 |
+
comparable scale (which is what the Elo vs. parameters chart needs); Elo weighs each win by the opponent's rating, so
|
| 65 |
+
beating a much weaker model earns almost nothing once ratings have settled. The match runs on the server
|
| 66 |
in the background: it finishes and counts even if the viewer leaves, and only one ranked match runs at a time
|
| 67 |
(pressing Play while one is running lets you watch it).
|
| 68 |
|
|
|
|
| 81 |
| `MAX_PARAMS` | `250000000` | Parameter limit |
|
| 82 |
| `MIN_PARAMS` | `50000` | Smallest model allowed (smaller ones are rejected and removed from the leaderboard) |
|
| 83 |
| `MAX_PIECES` | `500` | Piece cap per game |
|
| 84 |
+
| `MAX_PARAM_GAP` | `0` | Optional largest size difference between models picked for a ranked match (0 = no limit) |
|
| 85 |
+
| `LARGE_FROM` | `100000000` | Only with a gap: models this size or bigger can all play each other |
|
| 86 |
| `ALLOW_REMOTE_CODE` | `1` | Allow models with custom code (`trust_remote_code`) |
|
| 87 |
| `TORCH_THREADS` | `2` | CPU threads for inference |
|
| 88 |
|