And this is the part nobody is talking about:
Both models ship with 1M-token context. Both are within a few points of each other on most tasks. The gap between them is smaller than the gap between a good prompt and a bad one. The model is maybe 20% of your result. Your
Two models with 1M context, near equal, model just 20% of result
By
–