“On-Policy Self-Distillation for Reasoning Compression” This paper shows that the new bottleneck for reasoning LLMs isn’t “too little reasoning”, it’s that more reasoning tokens often increase mistakes. So they fixed it without RL length penalties or external verifiers, but all
On-Policy Self-Distillation Improves Reasoning LLM Efficiency
By
–
