Long uninterrupted runs are the new benchmark. A model that doesn't stop and start on a 50k-token refactor is shipping a different product than one that does, even if they score similar on short tasks. 4.6 stopping was masking how sensitive the previous loop was to noise.
Long-run model consistency becomes key performance benchmark
By
–