Benchmarking is especially hard when we don't even agree as to what the words involve mean. I believe that reasoning has been improved, but I am not sure what that actually translates to. The only way to figure out is to put in hours to test yourself?
Benchmarking AI Reasoning: Defining Metrics and Testing Improvements
By
–