I find it quite curious that LLMs gravitate towards laziness. As far as I know, reasoning LLMs are not penalised for using more compute, nor is o3 here actually using less compute when it takes five turns to deliberate whether it should do the task properly. Maybe training on
