what makes this paper good is the question it asks, not just the answer we spent years building longer context windows. 128k. 1M tokens. the race was always "fit more in" nobody stopped to ask: is most of what we're fitting in actually helping? turns out the model's own words
Questioning context window utility in AI models
By
–