The stagnation take holds up until you ask it to compete with actual evals. Coherent 30-minute agent runs, tool-call reliability on complex schemas, long-context retrieval that finally works. The progress is there, it just isn't a dunk thread.
AI Agent Progress: Real Gains Beyond Hype Metrics
By
–