botanic1-report / corpus.html
jean-livingmodels's picture
Link to the public Botanic1 collection (full slug) instead of the bare 'botanic1' path
0ff1e29 verified
Raw History Blame
13.9 kB
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Botanic1 pre-training corpus</title>
<script>(function(){try{var t=new URLSearchParams(location.search).get('__theme');if(t==='dark'||t==='light')document.documentElement.setAttribute('data-theme',t);}catch(e){}})();</script>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=Manrope:wght@500;600;700;800&family=Spectral:ital,wght@0,400;0,500;0,600;1,400&family=IBM+Plex+Mono:wght@400;500;600&display=swap">
<link rel="stylesheet" href="assets/site.css?v=20260908-use-case">
<script src="https://cdnjs.cloudflare.com/ajax/libs/d3/7.9.0/d3.min.js"></script>
</head>
<body>
<header class="topbar"><div class="inner"><a class="lockup" href="index.html"><div class="wordmark" aria-label="BOTANIC-1"><span class="cell">B</span><span class="cell">O</span><span class="cell">T</span><span class="cell">A</span><span class="cell">N</span><span class="cell">I</span><span class="cell">C</span><span class="cell tok" data-state="masked"><span class="probs"><i></i><i class="top"></i><i></i><i></i></span>?</span></div><span class="org"><svg class="living-models-mark" width="18" height="20" fill="currentColor" aria-hidden="true" focusable="false" xmlns="http://www.w3.org/2000/svg" version="1.1" viewBox="0 0 85.90 91.50"><rect x="55.1000" y="2.1000" width="3.2000" height="53.2000" /><rect x="55.1000" y="61.3000" width="3.2000" height="28.1000" /><rect x="68.9000" y="9.2000" width="3.2000" height="15.5000" /><rect x="68.9000" y="30.6000" width="3.2000" height="51.7000" /><rect x="82.7000" y="24.7000" width="3.2000" height="42.2000" /><rect x="41.4000" y="0.0000" width="3.2000" height="11.6000" /><rect x="41.4000" y="17.5000" width="3.2000" height="74.0000" /><rect x="27.6000" y="2.1000" width="3.2000" height="70.7000" /><rect x="27.6000" y="78.7000" width="3.2000" height="10.6000" /><rect x="13.8000" y="9.2000" width="3.2000" height="25.0000" /><rect x="13.8000" y="40.2000" width="3.2000" height="42.2000" /><rect x="0.0000" y="24.7000" width="3.2000" height="42.2000" /></svg>Living Models</span></a><nav aria-label="Pages"><a href="index.html">Overview</a><a href="factory.html">Model factory</a><a href="leaderboard.html">Leaderboard</a><a href="causal-variants.html">Causal variants</a><a href="sae-features.html">SAE features</a><a href="use-case.html">Use case</a><a href="corpus.html" aria-current="page">Corpus</a><a class="ext" href="https://huggingface.co/collections/living-models/botanic1-6a97f4e3c33f3d109a75057d" target="_blank" rel="noopener">Models</a></nav></div></header>
<div class="page">
<main class="prose">
<h1>A 320-genome corpus spanning the land plants</h1>
<p class="lede">Earlier studies all trained on fewer than 100 species (42 plant genomes for PlantBiMoE, 48 crop genomes for AgroNT and Botanic0, 65 for PlantCAD2), whereas the new Model Factory employs a 320-genome corpus spanning the land plants, thus yielding, to our knowledge, the broadest taxonomic coverage of any plant gLM to date.</p>
<p>The filters of the data pipeline leave 322 training and 4 held-out genomes (one per species); 320 of the training species contribute at least one window, for 7,207,506 windows and 59.0 Gbp of training sequence at 8,192 bp. Although the corpus spans 48 orders and 102 families, its coverage is strongly concentrated within angiosperms: 310 of the 320 species belong to the angiosperm crown group, with 241 eudicots and 61 monocots.</p>
<h2 id="sec-treemap">Windows per species</h2>
<p>Every rectangle is one species, sized by its number of 8,192 bp training windows and grouped by order. Colour is genome size, which the length weighting below deliberately keeps from dictating a species' share of the corpus.</p>
<div class="controls">
<div class="group"><input type="search" id="q" placeholder="Find a species, genus, family or order" aria-label="Find a species"></div>
<span class="count" id="count"></span>
</div>
<div class="treemap" id="treemap"></div>
<div class="legend"><span>Genome size</span><span class="ramp" id="ramp"></span><span id="ramp-min"></span><span>to</span><span id="ramp-max"></span><span style="margin-left:14px"><span class="swatch" style="background:var(--other)"></span>not resolved from released metadata</span></div>
<h2 id="weighting">Correcting for genome length</h2>
<p>The smallest genome in the corpus, <em>Arabidopsis thaliana</em> (0.12 Gbp), receives a weight of 1; the weight decreases with genome length until it reaches the floor of 0.5 at 0.48 Gbp, after which all larger genomes receive the same weight. Thus, for example, a 0.5 Gbp genome and the 11.9 Gbp genome of <em>Vicia faba</em> receive the same correction despite having widely different lengths. As a result, the correction reduces the tendency for larger genomes to contribute more windows but without erasing it completely. Empirically, doubling genome length increases a species' share of the 8 kbp windows by about 39% on average, compared with 48% without correction. Species with genomes larger than 2 Gbp therefore still occupy a larger median fraction of the corpus (0.50%) than species below 0.48 Gbp (0.19%), whereas perfectly uniform sampling across these 306 species would assign 0.33% to each.</p>
<h2 id="dedup">Near-duplicate windows</h2>
<p>Windows containing more than 20% <code>N</code> bases are removed. Near-duplicate windows are then filtered within each selected assembly, removing near-copies of the same window within species, but not orthologous sequence between species. Two windows are compared using the fraction of <em>k</em>-mers they have in common (Jaccard similarity) estimated with a per-window MinHash sketch (128 smallest hash values, each <em>k</em>-mer is canonicalised with its reverse complement so that strand orientation alone does not make two windows different). A window is discarded when its estimated Jaccard similarity to any window in the current set reaches 0.3. This removes roughly 15 to 25% of the sampled windows depending on the species pool and recipe; on the full 330-species ablation pool it removes 20.6% (65.6 to 52.1 Gbp).</p>
<h2 id="sec-orders">Orders</h2>
<div class="tablewrap"><table id="orders"></table></div>
<h2 id="ctx">The five context-extension datasets</h2>
<p>To expose the model to longer contexts, we create five datasets with window sizes 8,192, 16,384, 32,768, 65,536, and 131,072 bp. Those are produced by re-running our data pipeline at the target window length and increasing the contig-N50 floor accordingly. This limit rises faster than assembly contiguity improves, causing fewer species to qualify at each step. The 128 kbp dataset requires at least one long-read assembly per species, which is why it has the fewest species. Total base pairs per dataset range between 59 and 116 Gbp: fewer but longer windows offset the shrinking species pool.</p>
<div class="tablewrap">
<table>
<thead><tr><th>Window</th><th class="n">8 kbp</th><th class="n">16 kbp</th><th class="n">32 kbp</th><th class="n">64 kbp</th><th class="n">128 kbp</th></tr></thead>
<tbody>
<tr><td>Species qualifying</td><td class="n">326</td><td class="n">285</td><td class="n">273</td><td class="n">261</td><td class="n">238</td></tr>
<tr><td>Window sampling</td><td colspan="2" style="font-family:var(--sans);font-size:14px">coding-region-biased</td><td colspan="3" style="font-family:var(--sans);font-size:14px">repeat-content-biased</td></tr>
</tbody>
<caption>Longer windows (32, 64 and 128 kbp) almost always contain coding regions, so biasing by repeat content becomes more informative; this also lets more assemblies qualify even without precise genomic region annotations, which are no longer needed.</caption>
</table>
</div>
<p class="note">The 8 kbp corpus is released as the <a href="https://huggingface.co/datasets/living-models/Botanic1-pretraining">Botanic1-pretraining</a> dataset, with the original sequence case, augmentation margins, train/test separation, source manifests and assembly metadata.</p>
</main>
<footer class="site-footer">
<p>Living Models, Paris. Botanic1 is released for research use under the Living Models research licence.</p>
<p>Species and window counts from the corpus manifest; genome sizes from the report's assembly catalogue.</p>
</footer>
</div>
<div class="tip" id="tip" role="tooltip"></div>
<script src="assets/hero.js"></script>
<script>
(async function () {
const data = await (await fetch('data/corpus.json')).json();
const css = (v) => getComputedStyle(document.documentElement).getPropertyValue(v).trim();
const esc = (s) => String(s ?? '').replace(/&/g, '&amp;').replace(/</g, '&lt;');
const tip = document.getElementById('tip');
const showTip = (html, ev) => { tip.innerHTML = html; tip.style.display = 'block'; const pad = 14; let x = ev.clientX + pad, y = ev.clientY + pad; const r = tip.getBoundingClientRect(); if (x + r.width > innerWidth - 8) x = ev.clientX - r.width - pad; if (y + r.height > innerHeight - 8) y = ev.clientY - r.height - pad; tip.style.left = x + 'px'; tip.style.top = y + 'px'; };
const hideTip = () => { tip.style.display = 'none'; };
const total = data.totals.windows;
const gb = (bp) => bp == null ? null : (bp / 1e9);
const sizes = data.species.map(s => s.genome_bp).filter(Boolean);
const lo = d3.min(sizes), hi = d3.max(sizes);
const ramp = d3.scaleSequential(d3.interpolateLab(css('--paper-3'), css('--ours-deep'))).domain([Math.log10(lo), Math.log10(hi)]);
document.getElementById('ramp').style.background = `linear-gradient(90deg, ${css('--paper-3')}, ${css('--ours-deep')})`;
document.getElementById('ramp-min').textContent = gb(lo).toFixed(2) + ' Gbp';
document.getElementById('ramp-max').textContent = gb(hi).toFixed(1) + ' Gbp';
const root = d3.hierarchy(d3.group(data.species, d => d.order))
.sum(d => d.windows || 0)
.sort((a, b) => b.value - a.value);
const W = 1120, H = 720;
d3.treemap().size([W, H]).paddingOuter(3).paddingTop(20).paddingInner(1.5).round(true)(root);
const svg = d3.create('svg').attr('viewBox', `0 0 ${W} ${H}`).attr('role', 'img').attr('aria-label', 'Training windows per species, grouped by order');
const leaves = root.leaves();
const g = svg.append('g');
const cells = g.selectAll('g.cell').data(leaves).join('g').attr('class', 'cell').attr('transform', d => `translate(${d.x0},${d.y0})`);
cells.append('rect').attr('class', 'sp').attr('width', d => Math.max(0, d.x1 - d.x0)).attr('height', d => Math.max(0, d.y1 - d.y0))
.attr('fill', d => d.data.genome_bp ? ramp(Math.log10(d.data.genome_bp)) : css('--other'));
cells.filter(d => d.x1 - d.x0 > 64 && d.y1 - d.y0 > 16).append('text').attr('class', 'sp-label').attr('x', 4).attr('y', 12)
.text(d => { const w = d.x1 - d.x0; const s = d.data.species; return s.length * 6 > w - 8 ? s.split(' ')[0][0] + '. ' + s.split(' ').slice(1).join(' ') : s; })
.attr('fill', d => d.data.genome_bp && Math.log10(d.data.genome_bp) > (Math.log10(lo) + Math.log10(hi)) / 2 + 0.15 ? 'var(--paper)' : 'var(--ink)');
cells.on('mousemove', (ev, d) => showTip(`<b><i>${esc(d.data.species)}</i></b><br>${esc(d.data.order)} 路 ${esc(d.data.family)}<br>${d.data.windows.toLocaleString()} windows 路 ${(100 * d.data.windows / total).toFixed(2)}% of the corpus<br>genome ${d.data.genome_bp ? gb(d.data.genome_bp).toFixed(2) + ' Gbp' : 'not resolved'}`, ev)).on('mouseleave', hideTip);
// order frames and labels
const orders = root.children;
g.selectAll('rect.order').data(orders).join('rect').attr('class', 'order').attr('x', d => d.x0).attr('y', d => d.y0).attr('width', d => d.x1 - d.x0).attr('height', d => d.y1 - d.y0);
g.selectAll('text.order-label').data(orders.filter(d => d.x1 - d.x0 > 84)).join('text').attr('class', 'order-label').attr('x', d => d.x0 + 4).attr('y', d => d.y0 + 14)
.text(d => { const w = d.x1 - d.x0; const label = `${d.data[0]} 路 ${d.value.toLocaleString()}`; return label.length * 7 > w - 8 ? d.data[0] : label; });
document.getElementById('treemap').appendChild(svg.node());
document.getElementById('count').textContent = `${data.totals.species} species 路 ${total.toLocaleString()} windows 路 ${data.totals.orders} orders 路 ${data.totals.families} families`;
document.getElementById('q').addEventListener('input', (e) => {
const q = e.target.value.trim().toLowerCase();
let n = 0;
cells.select('rect.sp').classed('dim', d => { const hit = !q || [d.data.species, d.data.genus, d.data.family, d.data.order].some(v => v.toLowerCase().includes(q)); if (hit) n++; return !hit; });
document.getElementById('count').textContent = q ? `${n} of ${data.totals.species} species match` : `${data.totals.species} species 路 ${total.toLocaleString()} windows 路 ${data.totals.orders} orders 路 ${data.totals.families} families`;
});
// orders table
const byOrder = d3.rollups(data.species, v => ({ n: v.length, windows: d3.sum(v, d => d.windows), families: new Set(v.map(d => d.family)).size, top: v.slice().sort((a, b) => b.windows - a.windows)[0].species }), d => d.order)
.sort((a, b) => b[1].windows - a[1].windows);
let h = '<thead><tr><th>Order</th><th class="n">Species</th><th class="n">Families</th><th class="n">Windows</th><th class="n">Share</th><th>Largest contributor</th></tr></thead><tbody>';
byOrder.forEach(([o, v]) => { h += `<tr><td>${esc(o)}</td><td class="n">${v.n}</td><td class="n">${v.families}</td><td class="n">${v.windows.toLocaleString()}</td><td class="n">${(100 * v.windows / total).toFixed(1)}%</td><td><i>${esc(v.top)}</i></td></tr>`; });
h += '</tbody>';
document.getElementById('orders').innerHTML = h;
})();
</script>
</body>
</html>