OdysseyBench tests true agentic competence: • Contextual recall (across ~5K tokens)
• Multi-step reasoning
• Cross-app state tracking
• Intent inference through casual, noisy dialogue
• Tool execution (simulated, but grounded) It’s not toy data. It’s full workflow
OdysseyBench Evaluates Core AI Agentic Competencies
By
–