Today’s agents are “good” at: • Summarizing one doc
• Responding to one email
• Updating one file But in the real world? Tasks last 5 days, jump between apps, involve partial context, and require real memory. That’s what OdysseyBench simulates.
OdysseyBench Introduced to Evaluate Long-Running AI Agent Capabilities
By
–