# Reproduction: Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently ## Pages | Page | | --- | | [Executive summary](#/executive-summary) | | [Claim 1 — Golden-path SFT](#/claim-1-golden-path-sft) | | [Claim 2 — RLVR backtracking](#/claim-2-rlvr-backtracking) | | [Claim 3 — Inference separation](#/claim-3-inference-separation) | | [Claim 4 — Search agent](#/claim-4-search-agent) | | [Claim 5 — Trace distillation](#/claim-5-trace-distillation) | | [Methods](#/methods) | | [Negative controls](#/negative-controls) | | [Conclusion](#/conclusion) |