Beats flag-everything by 274.0, 95% CI [186,326]. Synthetic, not production-validated. Code: https://github.com/caiotheodoro/assay
Caio Theodoro
caiotheodoro
·
AI & ML interests
None yet
Recent Activity
updated a model 2 days ago
caiotheodoro/assay-challenger-grpo updated a Space 3 days ago
caiotheodoro/assay-demo published a Space 3 days ago
caiotheodoro/assay-demoOrganizations
None yet
Plumb: Ornith wrote the curriculum. Match won.
Self-proposed G702 training tasks lose alone, help as a supplement. Code: https://github.com/caiotheodoro/plumb
Suture: GPT-5.6 got 0.373. An 8B adapter got 0.959.
Binder vs issued policy, two page images in. 8B QLoRA, program oracle, 0.959 recall. Code: https://github.com/caiotheodoro/suture
LossBench: more accurate, 3.7x the loss.
qwen3.7-plus beats qwen3.6-plus on accuracy and loses 3.7x more. Code: https://github.com/caiotheodoro/lossbench
ReconForge: lost on accuracy. Caught every HIGH.
Loses accuracy to DeepSeek v4-flash, wins severity-weighted recall 0.901 vs 0.872 at 1.000 on HIGH. Code: https://github.com/caiotheodoro/reconforge
Assay: auditing RL environments, with error bars
Beats flag-everything by 274.0, 95% CI [186,326]. Synthetic, not production-validated. Code: https://github.com/caiotheodoro/assay
LossBench: more accurate, 3.7x the loss.
qwen3.7-plus beats qwen3.6-plus on accuracy and loses 3.7x more. Code: https://github.com/caiotheodoro/lossbench
Plumb: Ornith wrote the curriculum. Match won.
Self-proposed G702 training tasks lose alone, help as a supplement. Code: https://github.com/caiotheodoro/plumb
ReconForge: lost on accuracy. Caught every HIGH.
Loses accuracy to DeepSeek v4-flash, wins severity-weighted recall 0.901 vs 0.872 at 1.000 on HIGH. Code: https://github.com/caiotheodoro/reconforge
Suture: GPT-5.6 got 0.373. An 8B adapter got 0.959.
Binder vs issued policy, two page images in. 8B QLoRA, program oracle, 0.959 recall. Code: https://github.com/caiotheodoro/suture