Very interesting article from Anthropic about how, thanks to AI, they're able to streamline their process of researching and training better AI. Or as it's colloquially known: closing the loop
RESEARCH
-

Nemotron 3 Ultra open-weight release boasts impressive efficiency ratio
By
–

And another open-weight release. Nemotron 3 Ultra has an ultra impressive capability:efficiency ratio! Design-wise, it carries forward the Mamba-2-attention hybrid stack and LatentMoE introduced in the previous Super variant. But everything is a bit bigger.
-
AI’s core is context; solve context to solve any field.
By
–
AI at its core is a context problem. Solve context, and you can solve any field.
-
OpenAI and Anthropic predict faster self-recursive development
By
–
Yes. First OpenAI, now Anthropic. Both of them expect self recursive development to come faster than expected
-
Challenge of feeling AI acceleration despite real model improvements
By
–
A real problem with feeling the acceleration viscerally is that current models are really good and it is hard to feel the vibe difference on most individual tasks with new models, even as AIs continue to increase in ability by large amounts (which they actually are doing).
-
Anthropic’s serious push for recursive self-improvement
By
–
Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t
-

AutoLab: Encoding Persistence in Long-Horizon Agents
By
–
Outstanding paper on long-horizon agents. (bookmark it) Similar to humans, how do you make agents persist on a difficult task, and how is that useful? And which models today work well on this? This new work, AutoLab, explores this question and how encoding persistence in
-

Agent Arena evaluates model performance on real agentic tasks.
By
–
This is our most important eval yet – Agent Arena – it measures real performance of models on real agentic tasks. Our users use the Agent Arena, we monitor real signals (e.g. Bash Recovery) as well as their feedback on each task, without user knowing what model completed it.


