AI Dynamics

Global AI News Aggregator

About

On-Policy Distillation: Combining RL Error Correction with SFT Reward Density

On-policy distillation provides an elegant way to use the teacher model as a process reward model to provide dense reward while preventing SFT style "OOD shock" during rollout. Thinking Machines (@thinkymachines) Our latest post explores on-policy distillation, a training approach that unites the error-correcting relevance of RL with the reward density of SFT. When training it for math reasoning and as an internal chat assistant, we find that on-policy distillation can outperform other approaches for a fraction of the cost. thinkingmachines.ai/blog/on-… — https://nitter.net/thinkymachines/status/1982856272023302322#m

→ View original post on X — @lilianweng, 2025-10-27 17:31 UTC