AI Dynamics

Global AI News Aggregator

About

On-Policy Self-Distillation Improves Reasoning LLM Efficiency

“On-Policy Self-Distillation for Reasoning Compression” This paper shows that the new bottleneck for reasoning LLMs isn’t “too little reasoning”, it’s that more reasoning tokens often increase mistakes. So they fixed it without RL length penalties or external verifiers, but all

→ View original post on X — @askalphaxiv