Here's what it means for prompt engineering: They tracked individual answers across 32 reasoning budgets from 500 to 16,000 tokens. At ~7,000 tokens, something flips. The model starts abandoning correct answers MORE often than it finds new ones. They call it "negative flips."