This is our most important eval yet – Agent Arena – it measures real performance of models on real agentic tasks. Our users use the Agent Arena, we monitor real signals (e.g. Bash Recovery) as well as their feedback on each task, without user knowing what model completed it.
