Today’s agents are “good” at: • Summarizing one doc
• Responding to one email
• Updating one file But in the real world? Tasks last 5 days, jump between apps, involve partial context, and require real memory. That’s what OdysseyBench simulates.
LLMS
-
OdysseyBench Introduced to Evaluate Long-Running AI Agent Capabilities
By
–
-

OdysseyBench evaluates the real-world performance of LLM agents
By
–
This benchmark might kill your belief that LLM agents are “almost there.” OdysseyBench doesn’t test if an agent can summarize a file. It tests if it can survive a week at your office. And it exposes a brutal truth about today’s agents. Here’s what you need to know:
-
AI Model Explosion: New Releases From OpenAI Google Anthropic Grok
By
–
Choses qui n’existaient pas il y a un mois : – GPT-5
– GPT-5 mini
– GPT-5 nano
– GPT-OSS-120b
– GPT-OSS-20b
– Claude Opus 4.1
– Claude Sonnet 4.1
– Génie 3
– Google Deepthink advanced
– Google Gemma 270M
– Google Story book
– ChatGPT Agent
– Grok-4
– Grok Imagine
– Grok -
Custom AI Model Selection and System Prompts in Google Sheets
By
–
It's funny I had a script working in Google sheets doing this for about 18 months – the difference though is that I can select the models and a system prompt. One problem I have with these kinds of implementations is that they treat it as a generic 'AI', but I care a lot what
-
Reasoning Models Generate Significantly More Code Than Previous Models
By
–
The amount of code it was actually outputting was kind of crazy, especially comparing it to previous non-reasoning models, it is so much better than older ones
-
Gemini as a productivity multiplier for builders
By
–
Gemini feels like the ultimate productivity multiplier for any builder.
-
Grok AI accelerates application development cycles
By
–
App building guidance from Grok accelerates development cycles dramatically.
-

DLI2025 Attendees Master LLM and Reinforcement Learning in Hands-On Sessions
By
–



Following a day of deep-dive tutorials, attendees at #DLI2025 got to put their knowledge to the test in practical, hands-on sessions! @tejuafonja, Annie Qurat Ul Ain, and @jabez_magomere guided the Large Language Model practical, demystifying complex concepts and providing a chance to work with the technology directly. Simultaneously, Siddarth Singh, Sasha Abramowitz, and Ruan de Kock led the Reinforcement Learning session, turning theory into practice.
→ View original post on X — @shakir_za, 2025-08-19 06:38 UTC
-

Ai2 Launches MoNaCo: Multi-Step Evidence Synthesis Benchmark
By
–
Ai2 launched MoNaCo, a new eval that tests how well models stitch together evidence across dozens of sources It includes 1,315 multi‑step questions, retrieval, filtering & aggregation across text and tables, and 40+ distinct documents per query
-

AI Spreadsheet with Cell Agents Achieves General Availability
By
–
AI spreadsheet Paradigm, which has an agent in each cell, launched in general availability
— The Rundown AI (@TheRundownAI) 19 août 2025
It solves tasks that would normally require 30 hours of repetitive web research in just 30 seconds
Average throughput is 0.5B tokens/minpic.twitter.com/rsggugLoBJAI spreadsheet Paradigm, which has an agent in each cell, launched in general availability It solves tasks that would normally require 30 hours of repetitive web research in just 30 seconds Average throughput is 0.5B tokens/min