Upload folder using huggingface_hub
Browse files
README.md
CHANGED
|
@@ -24,7 +24,7 @@ AutoScientist Challenge (Part 2).
|
|
| 24 |
|
| 25 |
## Measured improvement
|
| 26 |
|
| 27 |
-
AutoScientist reported
|
| 28 |
|
| 29 |
Evaluation methodology, including the position-swap and dual-judge controls, is in
|
| 30 |
`EVAL.md` in the project repository. The held-out split used is published alongside the
|
|
|
|
| 24 |
|
| 25 |
## Measured improvement
|
| 26 |
|
| 27 |
+
AutoScientist reported best_win_rate = 0.4851 over 5 iterations against meta-llama/Llama-3.2-3B-Instruct. That is below the 0.50 break-even point, so this model does not improve on its baseline in English - see limitations.
|
| 28 |
|
| 29 |
Evaluation methodology, including the position-swap and dual-judge controls, is in
|
| 30 |
`EVAL.md` in the project repository. The held-out split used is published alongside the
|