jean-livingmodels commited on
Commit
be5ffad
·
verified ·
1 Parent(s): 77e0c61

Point the significance section at Table S7 and Figures 3d/3e of the report

Browse files
Files changed (1) hide show
  1. leaderboard.html +1 -1
leaderboard.html CHANGED
@@ -68,7 +68,7 @@
68
  <p class="chartnote">Vertical bars are 95% confidence intervals from the per-sample bootstrap. Tokens are converted to base pairs using approximations for 6-mer models. Models without a documented token count are omitted from the right panel.</p>
69
 
70
  <h2 id="uncertainty">Are leaderboard differences significant?</h2>
71
- <p>We assess leaderboard differences with two paired bootstrap tests, each using 20,000 replicates: one resampling the nine task families, and one resampling test examples within each benchmark cell. Together, they test whether observed margins are robust to variation across task families and to sampling noise in the benchmark. The technical report (v2) reports both tests: the table below reproduces the report's S<sub>bal</sub> uncertainty table, and the per-sample intervals are drawn as error bars and as paired differences to PlantCAD2-L in the report's leaderboard figure.</p>
72
  <p>The key result is that Botanic1-S performs on par with the much larger PlantCAD2-L: its score is slightly higher (+0.006), but the confidence intervals include zero, so we cannot resolve a significant difference between the two. In contrast, Botanic1-L and Botanic1-XL significantly outperform PlantCAD2-L under both tests. Botanic1-M (+0.012) is significant under the per-sample test but not under the family-level test, and shows a significant gain over Botanic1-S (+0.006) under both. Botanic1-S also significantly outperforms Carbon-8B and Evo&nbsp;2-7B under both tests.</p>
73
  <p>Overall, the results show that Botanic1 is already highly competitive at small scale, with larger variants delivering statistically supported gains over strong baselines.</p>
74
  <div class="tablewrap">
 
68
  <p class="chartnote">Vertical bars are 95% confidence intervals from the per-sample bootstrap. Tokens are converted to base pairs using approximations for 6-mer models. Models without a documented token count are omitted from the right panel.</p>
69
 
70
  <h2 id="uncertainty">Are leaderboard differences significant?</h2>
71
+ <p>We assess leaderboard differences with two paired bootstrap tests, each using 20,000 replicates: one resampling the nine task families, and one resampling test examples within each benchmark cell. Together, they test whether observed margins are robust to variation across task families and to sampling noise in the benchmark. The technical report (v2) reports both tests: the table below reproduces Supplementary Table S7, and the per-sample intervals are drawn as error bars in Figure 3d and as paired differences to PlantCAD2-L in Figure 3e.</p>
72
  <p>The key result is that Botanic1-S performs on par with the much larger PlantCAD2-L: its score is slightly higher (+0.006), but the confidence intervals include zero, so we cannot resolve a significant difference between the two. In contrast, Botanic1-L and Botanic1-XL significantly outperform PlantCAD2-L under both tests. Botanic1-M (+0.012) is significant under the per-sample test but not under the family-level test, and shows a significant gain over Botanic1-S (+0.006) under both. Botanic1-S also significantly outperforms Carbon-8B and Evo&nbsp;2-7B under both tests.</p>
73
  <p>Overall, the results show that Botanic1 is already highly competitive at small scale, with larger variants delivering statistically supported gains over strong baselines.</p>
74
  <div class="tablewrap">