Evaluation uses a mix of: • Exact/fuzzy matching on output files
• Execution-based checks (e.g., file created, event logged)
• Pass/fail based on fulfillment of all expected artifacts Every agent output is evaluated against the required plan, not just the final answer.
AGENTS
-
Evaluation Methodology for AI Agent Performance
By
–
-
OdysseyBench Evaluates Core AI Agentic Competencies
By
–
OdysseyBench tests true agentic competence: • Contextual recall (across ~5K tokens)
• Multi-step reasoning
• Cross-app state tracking
• Intent inference through casual, noisy dialogue
• Tool execution (simulated, but grounded) It’s not toy data. It’s full workflow -

OdysseyBench and HOMERAGENTS: A Multi-Agent Generation Pipeline
By
–
OdysseyBench is built using HOMERAGENTS, a multi-agent generation pipeline: • HOMERAGENTS+: Converts atomic tasks into multi-turn workflows • HOMERAGENTS-NEO: Synthesizes entirely new tasks using simulated app exploration Each task is: • Goal-driven
• Dialogue-based
• -

OdysseyBench: A New Benchmark for Long-Horizon AI Agent Tasks
By
–
What is OdysseyBench? A long-horizon benchmark of 602 multi-day, multi-application tasks: • Word, Excel, PDF, Email, Calendar
• 300 tasks from OfficeBench (restructured)
• 302 brand-new, complex tasks Each task is wrapped in realistic user–assistant dialogue over 5 days. -
OdysseyBench Introduced to Evaluate Long-Running AI Agent Capabilities
By
–
Today’s agents are “good” at: • Summarizing one doc
• Responding to one email
• Updating one file But in the real world? Tasks last 5 days, jump between apps, involve partial context, and require real memory. That’s what OdysseyBench simulates. -

OdysseyBench evaluates the real-world performance of LLM agents
By
–
This benchmark might kill your belief that LLM agents are “almost there.” OdysseyBench doesn’t test if an agent can summarize a file. It tests if it can survive a week at your office. And it exposes a brutal truth about today’s agents. Here’s what you need to know:
-

AI Spreadsheet with Cell Agents Achieves General Availability
By
–
AI spreadsheet Paradigm, which has an agent in each cell, launched in general availability
— The Rundown AI (@TheRundownAI) 19 août 2025
It solves tasks that would normally require 30 hours of repetitive web research in just 30 seconds
Average throughput is 0.5B tokens/minpic.twitter.com/rsggugLoBJAI spreadsheet Paradigm, which has an agent in each cell, launched in general availability It solves tasks that would normally require 30 hours of repetitive web research in just 30 seconds Average throughput is 0.5B tokens/min
-
Gods and Selves: The Reality of Belief and Consciousness
By
–
Gods are not necessarily less real than personal selfs. Your self may be a fiction of your brain, but it’s pretty real to you and those who believe in you
-
100+ Open Source AI Agents and RAG Tutorials Repository
By
–
I am that guy and here's the link to the GitHub Repo if you are wondering. I have spent last one year putting together 100+ opensource AI Agents and RAG tutorials. More here → @Saboo_Shubham_ https://
github.com/Shubhamsaboo/a
wesome-llm-apps
… -
100+ Free AI Agents and RAG Tutorials Released
By
–
Stay tuned for more such interesting posts → @Saboo_Shubham_ I have created 100+ AI Agents and RAG tutorials, 100% free and opensource. P.S: Don't forget to star the repo to show your support