The most useful finding for anyone using AI daily: When they capped reasoning at 60% of the model's natural length, it maintained 97% of peak accuracy. Longer natural outputs also correlated with lower accuracy. 71.9% accuracy under 4K tokens. 44.7% accuracy above 12K. The
Impact of Output Length on LLM Reasoning Accuracy
By
–