These micro-benchmarks are fun but I've found the model that "wins" changes depending on the exact framing of the prompt. Pattern matching != reasoning.
Micro-benchmarks Don’t Measure True Reasoning Capabilities
By
–
By
–
These micro-benchmarks are fun but I've found the model that "wins" changes depending on the exact framing of the prompt. Pattern matching != reasoning.