"Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?" Self-distillation can make LLMs look smarter by producing shorter, more confident reasoning traces, but in math it often takes out the model's uncertainty and self-correction signals. This can
