Outstanding paper on long-horizon agents. (bookmark it) Similar to humans, how do you make agents persist on a difficult task, and how is that useful? And which models today work well on this? This new work, AutoLab, explores this question and how encoding persistence in
AGENTS
-

Agent Arena evaluates model performance on real agentic tasks.
By
–
This is our most important eval yet – Agent Arena – it measures real performance of models on real agentic tasks. Our users use the Agent Arena, we monitor real signals (e.g. Bash Recovery) as well as their feedback on each task, without user knowing what model completed it.
-

Collaborative Battleship shows AI better at answering than asking questions
By
–
"Collaborative Battleship" game has revealed that AI agents are better at answering questions than asking them. CSAIL & SEAS had LMs play together, where Monte Carlo inference strategies helped small agents outpace the largest models at ~1% of the cost: https://
bit.ly/4afOE4T -
Building a bespoke harness for agent context: guide by Sydney Runkle
By
–
Agents are only as good as the context you give them. The job of a harness is to get the model the right context at the right time for a given task. Here’s a guide from @sydneyrunkle on how to build a bespoke harness for your use case.
-
Gary Marcus offers ‘truly impressed’ award for AI completing Baldur’s Gate 3
By
–
Offering a “damn, I’m truly impressed” award, for the first person or team to build a domain-general AI system that can play @baldursgate3, start to finish.
— Gary Marcus (@GaryMarcus) 4 juin 2026
Will make the prize especially sweet if you manage this decade. pic.twitter.com/M0f5BCh89pOffering a “damn, I’m truly impressed” award, for the first person or team to build a domain-general AI system that can play @baldursgate3
, start to finish. Will make the prize especially sweet if you manage this decade. -
Flows Agent lets you iterate and modify the pipeline dynamically
By
–
Flows Agent lets you iterate through conversation. Tell it to try a warmer voice, swap the background, or generate a version in Spanish. The agent modifies the pipeline and re-runs without rebuilding from scratch.
-
ElevenLabs introduces Flows Agent for automated creative workflow generation
By
–
Introducing Flows Agent in ElevenCreative.
— ElevenLabs (@ElevenLabs) 4 juin 2026
Describe what you want to create and the agent builds your entire workflow – selecting models, creating nodes, wiring connections, and running the generations. pic.twitter.com/m5Dyr77J15Introducing Flows Agent in ElevenCreative. Describe what you want to create and the agent builds your entire workflow – selecting models, creating nodes, wiring connections, and running the generations.
-
NVIDIA Announces Physical AI Agent Skills at CVPR2026
By
–
Building autonomous vehicles, robots, and vision AI takes more data than any team can collect on their own.
— NVIDIA (@nvidia) 4 juin 2026
At #CVPR2026, NVIDIA announced physical AI agent skills to help speed development: composable workflows that automate data generation, simulation, policy training and… pic.twitter.com/ZuHaTO683SBuilding autonomous vehicles, robots, and vision AI takes more data than any team can collect on their own. At #CVPR2026, NVIDIA announced physical AI agent skills to help speed development: composable workflows that automate data generation, simulation, policy training and
-

LangChain Supports NVIDIA Nemotron 3 Ultra with Deep Agents
By
–
LangChain supports @NVIDIA Nemotron 3 Ultra out of the box, with Day 0 support for Deep Agents. As a member of the Nemotron Coalition, we are excited to work with NVIDIA to make it even easier and more accessible to share and build on top of open models.
-

How a root agents.md ensures @agents .md is always created
By
–
in claude md just @agents
.md always i have a root agents md that tells agents when setting up new directories or repos to always do that.
