The assumption that "more thinking = better answers" was never tested until now. This paper tested it. The answer is no. For operators: if your AI output feels overengineered, bloated, or keeps second-guessing itself, the model isn't being thorough. It's losing confidence.