To test the broader usefulness of the AARs’ methods, we assessed how well they worked on two datasets the AARs hadn’t seen before. The AARs’ best-performing method successfully generalized to both coding and math tasks, though their second-best method only generalized to math.
AI
-
Claude Accelerates AI Alignment Research Experimentation Rate
By
–
AI models aren’t yet general-purpose alignment scientists. Progress isn't as easy to verify on most alignment research tasks: our AARs would find “fuzzier” research much harder. But our experiment does show that Claude can increase the rate of experimentation and exploration.
-

Automated Alignment Researchers Surpass Human Performance by 97%
By
–
Here, we measure success by the fraction of the “performance gap” we can close between the weak model and the potential of the strong model. After 7 days, human researchers closed it by 23%. Then, our Automated Alignment Researchers—Opus 4.6 with extra tools—closed it by 97%.
-
Anthropic Develops Automated Alignment Researcher with Claude
By
–
New Anthropic Fellows research: developing an Automated Alignment Researcher. We ran an experiment to learn whether Claude Opus 4.6 could accelerate research on a key alignment problem: using a weak AI model to supervise the training of a stronger one.
-
Autonomous Energy Systems Decision Speed Requirements
By
–
What's the fastest decision time your energy systems need to make autonomously? @IIoT_World @CRudinschi @agentic_factory @asokan_telecom @KADGLOBAL @Paul4innovating
-
Anthropic Introduces Routines for Claude Code in Research Preview
By
–
Anthropic released Routines for Claude Code in research preview.
— 🚨 AI News | TestingCatalog (@testingcatalog) 14 avril 2026
Routines are extended versions of Schedules, which can also be triggered via an API call or in response to an event.
Previously configured "Schedules" were also converted into Routines. https://t.co/gDrAqPrWHo pic.twitter.com/hzqHZMRE66Anthropic released Routines for Claude Code in research preview. Routines are extended versions of Schedules, which can also be triggered via an API call or in response to an event. Previously configured "Schedules" were also converted into Routines.
-
Democratic voting mechanisms prevent centralized AI policy decisions
By
–
Then have your country vote against the foreign policy plans. But be able to vote and not leave decision to one player alone
-
Multi-Agent System Optimizes CUDA Kernels for GPU Efficiency
By
–
The multi-agent system delivered optimizations that typically take experienced kernel engineers months or years. CUDA kernels are the core software supporting model training and inference. Faster kernels mean better GPU utilization and cheaper token costs.
-
Multi-Agent Architectures Excel Beyond Training Data Distribution
By
–
We see this as further validation that multi-agent architectures excel at novel problems outside training data distribution. These techniques will soon inform Cursor's core product.
-

AI System Optimizes Blackwell 200 GPUs Achieving 2x Speedups
By
–
The system learned to optimize Blackwell 200 GPUs from scratch, independently arriving at distinct optimization strategies across a long-tail of kernel problems. It outperformed baselines on 63% of problems and delivered more than 2x speedups on 19% of them.
