Spaces:
Running
Running
Point the significance section at Table S7 and Figures 3d/3e of the report
Browse files- leaderboard.html +1 -1
leaderboard.html
CHANGED
|
@@ -68,7 +68,7 @@
|
|
| 68 |
<p class="chartnote">Vertical bars are 95% confidence intervals from the per-sample bootstrap. Tokens are converted to base pairs using approximations for 6-mer models. Models without a documented token count are omitted from the right panel.</p>
|
| 69 |
|
| 70 |
<h2 id="uncertainty">Are leaderboard differences significant?</h2>
|
| 71 |
-
<p>We assess leaderboard differences with two paired bootstrap tests, each using 20,000 replicates: one resampling the nine task families, and one resampling test examples within each benchmark cell. Together, they test whether observed margins are robust to variation across task families and to sampling noise in the benchmark. The technical report (v2) reports both tests: the table below reproduces
|
| 72 |
<p>The key result is that Botanic1-S performs on par with the much larger PlantCAD2-L: its score is slightly higher (+0.006), but the confidence intervals include zero, so we cannot resolve a significant difference between the two. In contrast, Botanic1-L and Botanic1-XL significantly outperform PlantCAD2-L under both tests. Botanic1-M (+0.012) is significant under the per-sample test but not under the family-level test, and shows a significant gain over Botanic1-S (+0.006) under both. Botanic1-S also significantly outperforms Carbon-8B and Evo 2-7B under both tests.</p>
|
| 73 |
<p>Overall, the results show that Botanic1 is already highly competitive at small scale, with larger variants delivering statistically supported gains over strong baselines.</p>
|
| 74 |
<div class="tablewrap">
|
|
|
|
| 68 |
<p class="chartnote">Vertical bars are 95% confidence intervals from the per-sample bootstrap. Tokens are converted to base pairs using approximations for 6-mer models. Models without a documented token count are omitted from the right panel.</p>
|
| 69 |
|
| 70 |
<h2 id="uncertainty">Are leaderboard differences significant?</h2>
|
| 71 |
+
<p>We assess leaderboard differences with two paired bootstrap tests, each using 20,000 replicates: one resampling the nine task families, and one resampling test examples within each benchmark cell. Together, they test whether observed margins are robust to variation across task families and to sampling noise in the benchmark. The technical report (v2) reports both tests: the table below reproduces Supplementary Table S7, and the per-sample intervals are drawn as error bars in Figure 3d and as paired differences to PlantCAD2-L in Figure 3e.</p>
|
| 72 |
<p>The key result is that Botanic1-S performs on par with the much larger PlantCAD2-L: its score is slightly higher (+0.006), but the confidence intervals include zero, so we cannot resolve a significant difference between the two. In contrast, Botanic1-L and Botanic1-XL significantly outperform PlantCAD2-L under both tests. Botanic1-M (+0.012) is significant under the per-sample test but not under the family-level test, and shows a significant gain over Botanic1-S (+0.006) under both. Botanic1-S also significantly outperforms Carbon-8B and Evo 2-7B under both tests.</p>
|
| 73 |
<p>Overall, the results show that Botanic1 is already highly competitive at small scale, with larger variants delivering statistically supported gains over strong baselines.</p>
|
| 74 |
<div class="tablewrap">
|