Are AI benchmarks really measuring what matters for the economy? Enter Agents’ Last Exam (ALE) — a benchmark that tests AI agents on long, real-world professional tasks, not just puzzles. It covers 1,000+ tasks across 55 fields mapped to U.S. job classifications. The
Agents’ Last Exam: Benchmarking AI on real-world professional tasks for economy
By
–
