“On the Geometry of On-Policy Distillation” OPD is not just SFT mixed with RLVR. It has its own update geometry. This paper shows that OPD updates fewer weights than SFT and preserves pretrained structure better, while staying less constrained than RLVR. The key finding is
On-Policy Distillation geometry: fewer weight updates, preserves structure
By
–
