{ "seed": 915381, "steps": 200, "base": "/home/psikosen/canopy/checkpoints/canopy_browser_rlvr_sft.safetensors", "base_sha256": "2404b9e0c39f58286f82edb435a84dc1aba3af2e52f36c9d8692f855253fe8c0", "cases_sha256": "642c1fc1f1fd45523c17b6db412bfb57d4450ab49a9d8525842dddf3cd9be9ee", "data_sha256": "7134a6158aa90a86a749e29cc6a77eee6f4fb8c8de35662323db95598479b5d7", "training": "Binary and listwise semantic control scoring; 4 groups of 8 controls per update, final fixed-step checkpoint. Threshold/margin selected only on 96 validation groups.", "scope": "Eight known intents, new labels and randomized opaque identifiers, same controlled settings application. Scorer handles multi-button navigation menus only; existing SFT planner handles forms, backtracking, retries and reporting.", "comparison": "Same visible observations, 16-turn budget and private persisted-state grader. No private goal/label mapping or environment teacher in runtime. Missing and ambiguous cases require abstention without a write.", "pre_evaluation_amendment": "Before any evaluation: add a literal visible-description matching control; add eight separately frozen paraphrased-description stress cases to distinguish semantic transfer from exact string overlap. No training changes.", "stress_sha256": "1fa2345dba9d1aae57b7937467fc1f595a1b231348765f9fd8f417ca9850aa04", "evaluation_correction": "First run used separate ephemeral server ports per arm, which changed URL tokens. Archived under unmatched_urls. Final run shares one browser/server across all arms, clears state per episode, and asserts identical prompts/actions in literal-control fallback categories. All checkpoint weights, cases, threshold and margin unchanged. Missing/ambiguous success claims are graded as false even if a write happens to hit the nominal target." }