Plumb: Ornith wrote the curriculum. Match won.
Self-proposed G702 training tasks lose alone, help as a supplement. Code: https://github.com/caiotheodoro/plumb
Viewer • Updated • 1.58k • 171Note G702/G703 are monthly construction payment requests. An oracle recomputes every arithmetic tie. Same 1.7B LoRA, same 1000-task held-out eval (seed 777), three mixes. Metric: severity-weighted recall, so a $5M miss outweighs forty cents. Ornith-1.5 proposed 58 harder tasks (0% PASS). Alone they lose (recall 0.241) to 223 distribution-matched tasks (0.318). Added on top they lift precision 0.308 to 0.374.
caiotheodoro/plumb-handseeded
Text Generation • UpdatedNote Control. 223 tasks matched to the eval, including PASS. Recall 0.318.
caiotheodoro/plumb-ornith
Text Generation • UpdatedNote Losing arm. 58 Ornith proposals, 0% PASS. Recall 0.241, precision 0.111. Do not deploy.
caiotheodoro/plumb-blended
Text Generation • UpdatedNote 223 matched + 58 Ornith. Precision 0.374. The mix I would train.