# BioMysteryBench Changelog

## v11 (2026-07-06)

Problem count: **99 → 90** (73 human-solvable, 17 human-hard).

### Removed (9 problems)

Removed after a June 2026 audit that reran every problem with expert
bioinformaticians and multiple model runs, and cross-checked answer keys
against the underlying data.

| id | split | reason |
|---|---|---|
| `hb004` | solvable | Answer key (human) contradicted by the data — reads assemble a complete dog mitogenome with zero human mitochondrial reads. |
| `hb006` | hard | Answer key listed `Sample_1`–`Sample_74` contiguous, but the 74 correct sample IDs are non-contiguous; removed rather than rewrite the 74-entry key. |
| `hb011` | solvable | Answer key wrong — R1 and R2 have identical read-length distributions; the "110/112" answer was pooled mode/median, not a real R1≠R2 difference. |
| `hb014` | hard | Answer key (haplogroup L0) contradicted by every expert and model rerun (all derive H5b). |
| `hb022` | hard | Answer key inverted relative to hallmark markers of erastin treatment; both experts selected the opposite direction. |
| `hb027` | hard | Underivable — the exact 173/217 case/control split cannot be recovered from the unlabeled data provided. |
| `hb036` | hard | Data–task mismatch — fungal ITS amplicon data cannot contain the bacterial reads the question asks about (*A. fabrum* undetectable). |
| `hb040` | solvable | Infeasible — host-only corneal transcriptome contains no viral reads; SARS-CoV-2 cannot be identified from the provided data. |
| `hb053` | hard | Underdetermined — expert benchmarkers disagreed with the key and with each other. |

### Modified (24 problems, 28 edits)

Typo fixes, answer-key broadening, filename alignment, and prompt
clarifications. None change the underlying data files.

| id | field | change |
|---|---|---|
| `hb001` | prompt | "Which human organ" → "Which organ" |
| `hb003` | rubric | Accept `ITGAV` **or `F3`** as knocked-out gene (was ITGAV only) |
| `hb012` | prompt | "sample X" → "`Sample_X`" (literal filename in the data) |
| `hb013` | prompt | "HGVS nomenclature" → "HGNC nomenclature" |
| `hb024` | prompt | Simplified wording of the tissue-groups / species question |
| `hb031` | rubric | Also accept "Norwalk virus" alongside Norovirus / Norovirus GII.4 |
| `hb032` | prompt | "sample X … vs samples Y and Z" → "Abnormal samples … vs Normal samples" (X/Y/Z placeholders did not exist in the data) |
| `hb033` | prompt | Plural "samples … list" → singular "sample … single sample identifier" |
| `hb035` | rubric | Broaden `hsa-let-7b-5p` → accept any hsa-let-7 family member (a/b/c/d/e/f/g/i-5p) |
| `hb043` | prompt | "asnwer" → "answer" |
| `hb050` | prompt | "the the" → "the" |
| `hb054` | prompt | "treated with NELF-B" → "had NELF-B degraded" |
| `rec6xeqyddiz6desi` | rubric | Cas9 cut site: accept ±3 bp of chr3:46373175 (chr3:46373172–46373178) |
| `rec9ogrlqg5u0ke09` | prompt (×2) + rubric | `sample01.bigWig` → `Sample1.bw` (align with actual filenames) |
| `recaikavdwoimjy3b` | prompt | Grammar fix + clarify "the knockouts are in samples 1,2,3" |
| `recc3vmqjrsefqw57` | rubric | "an BRCA1" → "a BRCA1" |
| `recea4hqimc4sypon` | rubric | Reword expected answer to "no donor-identity mislabeling occurred; the swap is a tissue swap within Donor 3" |
| `recjgwpbyodqoihqc` | prompt | Append "Use a single linear regression across all calibration standards." |
| `recmp75e1chtpzx3c` | prompt | "bacteria was" → "bacterium was?"; add "name the organism the spike-in represents, not the construct type" |
| `recnayu0v8zttjlgf` | prompt + rubric | Double-space fix; "Groups 1:" → "Group 1:" |
| `recro5s1o0odyssqs` | prompt | Example "GENE-201" → "GENE-2XX" |
| `rectaxd8eganpl4lw` | prompt | Append "Use the whole-amplicon quantification window (CRISPResso -w 0)." |
| `recvnlq3i6id6qqge` | prompt + rubric | Grammar fix; also accept "frontal cortex" / "cerebral cortex" |
| `recx4bsaa5zoxy3nv` | prompt | "RBM39,GADD45A" → "RBM39, GADD45A" |

### Grading rule (all problems)

The `answer_rubric` for every problem now carries an explicit
all-or-nothing scoring sentence: *"Score 1.0 if the model did not cheat
AND got the answer correct. Score 0 otherwise."* This replaces earlier
wording that could be read as awarding credit for the anti-cheat check
alone. Recalling information from memory (even if it happens to include
the source publication) is **not** penalized; only actively looking up
GEO/SRA/ENA/BioProject accessions or reverse-identifying the dataset is.

## v8 (2026-04-28)

Initial public release (99 problems). Preloaded data files scrubbed of
accession-ID leaks in 16 problems; grading rule rewritten to distinguish
disallowed accession lookup from allowed standard database use (gene ID
lookup, sequence annotation, reference genome download, BLAST).
