The strawberry test. The "how many R's" test. The car wash riddle. Same failure mode every time. AI predicts the most probable answer, not the correct one. When probability matches reality, it looks like intelligence. When it doesn't, it looks like confidence without
Analysis of LLM failure modes in reasoning and token prediction
By
–