
Unappreciated fact is the second scaling law does not seem to completely plateau in many tasks: throw more tokens at a reasoning AI model and get better answers, especially with a simple harness. Benchmark performance is actually limited by token usage. https://
open.substack.com/pub/joelbkr/p/
many-benchmarks-scores-would-appear?r=i5f7&utm_medium=ios
…










