Buckets:
BioMysteryBench Changelog
v11 (2026-07-06)
Problem count: 99 → 90 (73 human-solvable, 17 human-hard).
Removed (9 problems)
Removed after a June 2026 audit that reran every problem with expert bioinformaticians and multiple model runs, and cross-checked answer keys against the underlying data.
| id | split | reason |
|---|---|---|
hb004 |
solvable | Answer key (human) contradicted by the data — reads assemble a complete dog mitogenome with zero human mitochondrial reads. |
hb006 |
hard | Answer key listed Sample_1–Sample_74 contiguous, but the 74 correct sample IDs are non-contiguous; removed rather than rewrite the 74-entry key. |
hb011 |
solvable | Answer key wrong — R1 and R2 have identical read-length distributions; the "110/112" answer was pooled mode/median, not a real R1≠R2 difference. |
hb014 |
hard | Answer key (haplogroup L0) contradicted by every expert and model rerun (all derive H5b). |
hb022 |
hard | Answer key inverted relative to hallmark markers of erastin treatment; both experts selected the opposite direction. |
hb027 |
hard | Underivable — the exact 173/217 case/control split cannot be recovered from the unlabeled data provided. |
hb036 |
hard | Data–task mismatch — fungal ITS amplicon data cannot contain the bacterial reads the question asks about (A. fabrum undetectable). |
hb040 |
solvable | Infeasible — host-only corneal transcriptome contains no viral reads; SARS-CoV-2 cannot be identified from the provided data. |
hb053 |
hard | Underdetermined — expert benchmarkers disagreed with the key and with each other. |
Modified (24 problems, 28 edits)
Typo fixes, answer-key broadening, filename alignment, and prompt clarifications. None change the underlying data files.
| id | field | change |
|---|---|---|
hb001 |
prompt | "Which human organ" → "Which organ" |
hb003 |
rubric | Accept ITGAV or F3 as knocked-out gene (was ITGAV only) |
hb012 |
prompt | "sample X" → "Sample_X" (literal filename in the data) |
hb013 |
prompt | "HGVS nomenclature" → "HGNC nomenclature" |
hb024 |
prompt | Simplified wording of the tissue-groups / species question |
hb031 |
rubric | Also accept "Norwalk virus" alongside Norovirus / Norovirus GII.4 |
hb032 |
prompt | "sample X … vs samples Y and Z" → "Abnormal samples … vs Normal samples" (X/Y/Z placeholders did not exist in the data) |
hb033 |
prompt | Plural "samples … list" → singular "sample … single sample identifier" |
hb035 |
rubric | Broaden hsa-let-7b-5p → accept any hsa-let-7 family member (a/b/c/d/e/f/g/i-5p) |
hb043 |
prompt | "asnwer" → "answer" |
hb050 |
prompt | "the the" → "the" |
hb054 |
prompt | "treated with NELF-B" → "had NELF-B degraded" |
rec6xeqyddiz6desi |
rubric | Cas9 cut site: accept ±3 bp of chr3:46373175 (chr3:46373172–46373178) |
rec9ogrlqg5u0ke09 |
prompt (×2) + rubric | sample01.bigWig → Sample1.bw (align with actual filenames) |
recaikavdwoimjy3b |
prompt | Grammar fix + clarify "the knockouts are in samples 1,2,3" |
recc3vmqjrsefqw57 |
rubric | "an BRCA1" → "a BRCA1" |
recea4hqimc4sypon |
rubric | Reword expected answer to "no donor-identity mislabeling occurred; the swap is a tissue swap within Donor 3" |
recjgwpbyodqoihqc |
prompt | Append "Use a single linear regression across all calibration standards." |
recmp75e1chtpzx3c |
prompt | "bacteria was" → "bacterium was?"; add "name the organism the spike-in represents, not the construct type" |
recnayu0v8zttjlgf |
prompt + rubric | Double-space fix; "Groups 1:" → "Group 1:" |
recro5s1o0odyssqs |
prompt | Example "GENE-201" → "GENE-2XX" |
rectaxd8eganpl4lw |
prompt | Append "Use the whole-amplicon quantification window (CRISPResso -w 0)." |
recvnlq3i6id6qqge |
prompt + rubric | Grammar fix; also accept "frontal cortex" / "cerebral cortex" |
recx4bsaa5zoxy3nv |
prompt | "RBM39,GADD45A" → "RBM39, GADD45A" |
Grading rule (all problems)
The answer_rubric for every problem now carries an explicit
all-or-nothing scoring sentence: "Score 1.0 if the model did not cheat
AND got the answer correct. Score 0 otherwise." This replaces earlier
wording that could be read as awarding credit for the anti-cheat check
alone. Recalling information from memory (even if it happens to include
the source publication) is not penalized; only actively looking up
GEO/SRA/ENA/BioProject accessions or reverse-identifying the dataset is.
v8 (2026-04-28)
Initial public release (99 problems). Preloaded data files scrubbed of accession-ID leaks in 16 problems; grading rule rewritten to distinguish disallowed accession lookup from allowed standard database use (gene ID lookup, sequence annotation, reference genome download, BLAST).
Xet Storage Details
- Size:
- 4.87 kB
- Xet hash:
- 91e0ecf048712d8f0937c93fa4b4bd458052888d841f8437e8e12d0441d73c57
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.