AI research is a series of next-step decisions. We looked at sessions where a human researcher took a wrong turn, showed Claude the session up to that point, and asked it what to do next. Mythos Preview improved on humans 64% of the time—up from 22% in 2024.
RESEARCH
-
AI Self-Improvement Plausible if Trends Continue, But Research Judgment Lacks
By
–
None of this guarantees recursive self-improvement is on the horizon. It’s not yet clear that Claude is capable of research judgment—of choosing the right problems to work on. But if these trends continue, AI systems designing and building their own successors is plausible. This
-
Claude Accelerating AI Development: Recursive Self-Improvement Faster Than Expected
By
–
Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention.
-

Collaborative Battleship shows AI better at answering than asking questions
By
–
"Collaborative Battleship" game has revealed that AI agents are better at answering questions than asking them. CSAIL & SEAS had LMs play together, where Monte Carlo inference strategies helped small agents outpace the largest models at ~1% of the cost: https://
bit.ly/4afOE4T -
Gary Marcus offers ‘truly impressed’ award for AI completing Baldur’s Gate 3
By
–
Offering a “damn, I’m truly impressed” award, for the first person or team to build a domain-general AI system that can play @baldursgate3, start to finish.
— Gary Marcus (@GaryMarcus) 4 juin 2026
Will make the prize especially sweet if you manage this decade. pic.twitter.com/M0f5BCh89pOffering a “damn, I’m truly impressed” award, for the first person or team to build a domain-general AI system that can play @baldursgate3
, start to finish. Will make the prize especially sweet if you manage this decade. -
Using LLMs to substantiate private allegations with public evidence
By
–
I have a talk at Manifest about ParaLLM Construction coming up, about recent reporting I did which used LLMs to find public evidence substantiating (and letting me use) private allegations.
-
NVIDIA Announces Physical AI Agent Skills at CVPR2026
By
–
Building autonomous vehicles, robots, and vision AI takes more data than any team can collect on their own.
— NVIDIA (@nvidia) 4 juin 2026
At #CVPR2026, NVIDIA announced physical AI agent skills to help speed development: composable workflows that automate data generation, simulation, policy training and… pic.twitter.com/ZuHaTO683SBuilding autonomous vehicles, robots, and vision AI takes more data than any team can collect on their own. At #CVPR2026, NVIDIA announced physical AI agent skills to help speed development: composable workflows that automate data generation, simulation, policy training and
-
Paper link to NVIDIA Nemotron-3 Ultra Technical Report
By
–
4/ Paper: https://
research.nvidia.com/labs/nemotron/
files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf
… -

NVIDIA Nemotron 3 Ultra: open 550B hybrid Mamba-Attention MoE
By
–


1/ NVIDIA shipped Nemotron 3 Ultra today, a fully open 550B model with 55B active params, with the weights, training data, and complete recipe all released openly. That alone is rare at this scale. The headline however actually is speed. Ultra is a hybrid Mamba-Attention MoE, an
-

AI consciousness: The pressing philosophical question of our age, debated on Hackernews
By
–
Intense discussion at the moment on Hackernews, reacting to Ted Chiang's article. "Can AI be conscious" is a pressing question, and could become the prime philosophical question of our age.
