shubhxho commited on
Commit
5c1f878
·
verified ·
1 Parent(s): 2874883

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +29 -11
README.md CHANGED
@@ -59,12 +59,16 @@ model-index:
59
  name: Chess play, absolute rating anchor
60
  dataset:
61
  type: stockfish-uci-elo
62
- name: Stockfish with UCI_LimitStrength, four settings
63
  metrics:
64
  - type: elo
65
- name: Implied Elo on Stockfish's UCI_Elo scale
66
  value: 2800
67
- args: 200 games per setting at 100ms/move; anchors spread 2741-2881, so treat as +/-70
 
 
 
 
68
  verified: false
69
  - task:
70
  type: other
@@ -238,14 +242,28 @@ For an absolute figure, the engine was played against Stockfish under
238
 
239
  | Stockfish `UCI_Elo` | score | implied |
240
  |---|---|---|
241
- | 2200 | 0.958 | 2741 |
242
- | 2500 | 0.853 | 2805 |
243
- | 2800 | 0.465 | **2776** |
244
- | 3000 | 0.335 | 2881 |
245
-
246
- Call it **2800**. The 2800 row deserves the most weight because it is nearest
247
- parity and extrapolates least. The four anchors disagree by 140 Elo, and that
248
- spread is the honest precision — this locates the engine on someone else's scale
 
 
 
 
 
 
 
 
 
 
 
 
 
 
249
  rather than rating it, and it is not a CCRL or FIDE number.
250
 
251
  ## The output gain
 
59
  name: Chess play, absolute rating anchor
60
  dataset:
61
  type: stockfish-uci-elo
62
+ name: Stockfish with UCI_LimitStrength, five settings from 2600 to 3000
63
  metrics:
64
  - type: elo
65
+ name: Implied Elo on Stockfish's UCI_Elo scale (0.5 crossover)
66
  value: 2800
67
+ args: >-
68
+ 1500 games, 300 at each of five settings, 100ms/move. Maximum-likelihood
69
+ fit 2819 +/- 19, crossover interpolation 2783; quote as +/-40. Stockfish's
70
+ nominal scale measures compressed here (fitted slope 0.83), so the crossover
71
+ is the slope-independent estimate.
72
  verified: false
73
  - task:
74
  type: other
 
242
 
243
  | Stockfish `UCI_Elo` | score | implied |
244
  |---|---|---|
245
+ | 2600 | 0.772 | 2812 |
246
+ | 2700 | 0.638 | 2799 |
247
+ | 2800 | 0.472 | 2780 |
248
+ | 2900 | 0.410 | 2837 |
249
+ | 3000 | 0.328 | 2876 |
250
+
251
+ Call it **2800**, and mean it loosely. A maximum-likelihood fit over all 1500
252
+ games says 2819 ± 19; the point where the score actually crosses 0.5 says 2783.
253
+ Quote the range, not either end — ±40 is honest, ±19 is not.
254
+
255
+ They disagree for a reason worth knowing if you ever calibrate anything this
256
+ way. The one-parameter fit leaves residuals that drift monotonically with the
257
+ setting (-0.008, -0.027, -0.056, +0.024, +0.067), meaning this engine loses less
258
+ to Stockfish's strongest settings than the logistic model predicts. Letting the
259
+ slope float fits it at 0.83 — a hundred of Stockfish's nominal points behave like
260
+ roughly eighty-three real ones across this range. `UCI_LimitStrength` hits its
261
+ target by degrading play in discrete internal steps, so its scale has no
262
+ particular reason to be linear, and measured here it isn't.
263
+
264
+ That is why the 0.5 crossover is the defensible number: two engines scoring 0.5
265
+ against each other are equal by definition, and that point doesn't depend on the
266
+ slope being correct. This locates the engine on someone else's approximate scale
267
  rather than rating it, and it is not a CCRL or FIDE number.
268
 
269
  ## The output gain