dna-benchmark / README.md
Pedro Dias
changed color
83f7a75
|
Raw History Blame Contribute Delete
1.44 kB
---
title: DNA Benchmark
emoji: 🥇
colorFrom: red
colorTo: indigo
sdk: gradio
app_file: app.py
pinned: true
license: apache-2.0
short_description: DNA Foundational Models leaderboard
sdk_version: 5.19.0
---
# DNA Benchmark Leaderboard
Interactive leaderboard comparing DNA foundation models across genomics classification tasks.
Data is loaded from [`lokahq/genomic-benchmark-metrics`](https://huggingface.co/datasets/lokahq/genomic-benchmark-metrics) on startup.
## Development
```bash
# Install uv (if not already installed)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install dependencies
uv sync
# Run locally
make app
# Format code
make style
```
## Context length bubble chart
The bubble chart requires mean sequence lengths per task to be pre-computed once:
```bash
make seq-lengths # computes lengths and writes tools/task_seq_lengths.json
git add tools/task_seq_lengths.json && git commit -m "populate task seq lengths"
```
The chart is a no-op until `tools/task_seq_lengths.json` is populated and committed.
## Configuration
| File | Purpose |
|---|---|
| `src/constants.py` | Task categories, metric names, seq length lookup |
| `src/data.py` | HF Hub loading, deduplication, filtering |
| `src/plots.py` | All Plotly figure factories |
| `tools/task_seq_lengths.json` | Mean sequence lengths per task (committed) |
| `.env` | `HF_TOKEN` and `LEADERBOARD_DATASET` (use HF Spaces secrets in deployment) |