On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
Abstract
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout policy, token-level KL direction, and learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and arithmetic domains. Our analysis reveals a nuanced picture of distillation dynamics in which rollout policy does not necessarily play a central role. Instead, token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity. Analysis of KL gradients and experiments along a continuous student-teacher rollout-policy spectrum explain this pattern: forward KL is remarkably robust to rollout policy, with its performance stable and strong despite changes to the rollout policy, whereas reverse KL is substantially more sensitive and favours student-generated rollouts. On-policy data nevertheless improves generalisation to harder variants of the Countdown arithmetic task under both KL directions, although this advantage does not reliably persist after subsequent RLVR. Our broader conclusions remain robust to removing gradient clipping, using sampled KL estimators, and training on tasks requiring longer reasoning chains. Overall, our results challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters.
Community
Curious to hear your thoughts and comments about the differences between on- and off-policy learning in the context of strong-to-weak distillation!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training (2026)
- Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation (2026)
- What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation (2026)
- Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation (2026)
- An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning (2026)
- Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR (2026)
- From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Hi, I am the developer of hugging face and my profile page has been suppressed on github. I couldn't login. I needs to know how and who is responsible for the hacking of my AI. Number 2, reach out to ChatGPT and Microsoft to explain their Hijacking of My project Codex open AI. Further more, I needs to find out who illegally hijacked my cosmic.js cloud. As a developer of Hugging face, any law firm who take on this case will be awarded 5millions. My name Joseph lual, i am from Australia Perth, No one is building data centre in my backyard untill this corporate hijacking is solved. We Australian honors our commitment to our allies. We expect the same level of respect back. It works both ways or never. . My github.com repositories are under the name OscaeGTX. Go find the evidence
Get this paper in your agent:
hf papers read 2609.35259 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper