It gets worse for simple questions. Easy problems hit the overthinking zone at 2,000 tokens. Hard problems don't hit it until 8,000. Translation: the simpler the question, the faster the model starts hurting its own answer by thinking longer. Optimal reasoning length varies
Analysis of LLM reasoning performance relative to token length
By
–