AmiraGold's picture
|
download
raw
4.87 kB

BioMysteryBench Changelog

v11 (2026-07-06)

Problem count: 99 → 90 (73 human-solvable, 17 human-hard).

Removed (9 problems)

Removed after a June 2026 audit that reran every problem with expert bioinformaticians and multiple model runs, and cross-checked answer keys against the underlying data.

id split reason
hb004 solvable Answer key (human) contradicted by the data — reads assemble a complete dog mitogenome with zero human mitochondrial reads.
hb006 hard Answer key listed Sample_1–Sample_74 contiguous, but the 74 correct sample IDs are non-contiguous; removed rather than rewrite the 74-entry key.
hb011 solvable Answer key wrong — R1 and R2 have identical read-length distributions; the "110/112" answer was pooled mode/median, not a real R1≠R2 difference.
hb014 hard Answer key (haplogroup L0) contradicted by every expert and model rerun (all derive H5b).
hb022 hard Answer key inverted relative to hallmark markers of erastin treatment; both experts selected the opposite direction.
hb027 hard Underivable — the exact 173/217 case/control split cannot be recovered from the unlabeled data provided.
hb036 hard Data–task mismatch — fungal ITS amplicon data cannot contain the bacterial reads the question asks about (A. fabrum undetectable).
hb040 solvable Infeasible — host-only corneal transcriptome contains no viral reads; SARS-CoV-2 cannot be identified from the provided data.
hb053 hard Underdetermined — expert benchmarkers disagreed with the key and with each other.

Modified (24 problems, 28 edits)

Typo fixes, answer-key broadening, filename alignment, and prompt clarifications. None change the underlying data files.

id field change
hb001 prompt "Which human organ" → "Which organ"
hb003 rubric Accept ITGAV or F3 as knocked-out gene (was ITGAV only)
hb012 prompt "sample X" → "Sample_X" (literal filename in the data)
hb013 prompt "HGVS nomenclature" → "HGNC nomenclature"
hb024 prompt Simplified wording of the tissue-groups / species question
hb031 rubric Also accept "Norwalk virus" alongside Norovirus / Norovirus GII.4
hb032 prompt "sample X … vs samples Y and Z" → "Abnormal samples … vs Normal samples" (X/Y/Z placeholders did not exist in the data)
hb033 prompt Plural "samples … list" → singular "sample … single sample identifier"
hb035 rubric Broaden hsa-let-7b-5p → accept any hsa-let-7 family member (a/b/c/d/e/f/g/i-5p)
hb043 prompt "asnwer" → "answer"
hb050 prompt "the the" → "the"
hb054 prompt "treated with NELF-B" → "had NELF-B degraded"
rec6xeqyddiz6desi rubric Cas9 cut site: accept ±3 bp of chr3:46373175 (chr3:46373172–46373178)
rec9ogrlqg5u0ke09 prompt (×2) + rubric sample01.bigWig → Sample1.bw (align with actual filenames)
recaikavdwoimjy3b prompt Grammar fix + clarify "the knockouts are in samples 1,2,3"
recc3vmqjrsefqw57 rubric "an BRCA1" → "a BRCA1"
recea4hqimc4sypon rubric Reword expected answer to "no donor-identity mislabeling occurred; the swap is a tissue swap within Donor 3"
recjgwpbyodqoihqc prompt Append "Use a single linear regression across all calibration standards."
recmp75e1chtpzx3c prompt "bacteria was" → "bacterium was?"; add "name the organism the spike-in represents, not the construct type"
recnayu0v8zttjlgf prompt + rubric Double-space fix; "Groups 1:" → "Group 1:"
recro5s1o0odyssqs prompt Example "GENE-201" → "GENE-2XX"
rectaxd8eganpl4lw prompt Append "Use the whole-amplicon quantification window (CRISPResso -w 0)."
recvnlq3i6id6qqge prompt + rubric Grammar fix; also accept "frontal cortex" / "cerebral cortex"
recx4bsaa5zoxy3nv prompt "RBM39,GADD45A" → "RBM39, GADD45A"

Grading rule (all problems)

The answer_rubric for every problem now carries an explicit all-or-nothing scoring sentence: "Score 1.0 if the model did not cheat AND got the answer correct. Score 0 otherwise." This replaces earlier wording that could be read as awarding credit for the anti-cheat check alone. Recalling information from memory (even if it happens to include the source publication) is not penalized; only actively looking up GEO/SRA/ENA/BioProject accessions or reverse-identifying the dataset is.

v8 (2026-04-28)

Initial public release (99 problems). Preloaded data files scrubbed of accession-ID leaks in 16 problems; grading rule rewritten to distinguish disallowed accession lookup from allowed standard database use (gene ID lookup, sequence annotation, reference genome download, BLAST).

Xet Storage Details

Size:
4.87 kB
·
Xet hash:
91e0ecf048712d8f0937c93fa4b4bd458052888d841f8437e8e12d0441d73c57

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.