The speedup isn’t just in volume. On open-ended coding problems where answers are unclear, Claude’s success rate is now 76%—a 50 point jump in just 6 months. Many engineers also say Claude’s code quality is now on par with human code; we expect it to be better within the year.
MACHINE LEARNING
-
New memory system enables review and control of ChatGPT’s context
By
–
With the new memory system, you can review and steer what ChatGPT remembers through a memory summary, with more visibility and control over how context is used. pic.twitter.com/kXMAds0g3q
— OpenAI (@OpenAI) 4 juin 2026With the new memory system, you can review and steer what ChatGPT remembers through a memory summary, with more visibility and control over how context is used.
-

Collaborative Battleship shows AI better at answering than asking questions
By
–
"Collaborative Battleship" game has revealed that AI agents are better at answering questions than asking them. CSAIL & SEAS had LMs play together, where Monte Carlo inference strategies helped small agents outpace the largest models at ~1% of the cost: https://
bit.ly/4afOE4T -
ElevenCreative Flows links 50+ models with voice, music, SFX
By
–
ElevenCreative Flows connects 50+ image and video models with voice, music, and SFX on one canvas. Creators and marketers use it to chain modalities together into complete pipelines and test creative variants across products, languages, and formats.
-
ElevenLabs introduces Flows Agent for automated creative workflow generation
By
–
Introducing Flows Agent in ElevenCreative.
— ElevenLabs (@ElevenLabs) 4 juin 2026
Describe what you want to create and the agent builds your entire workflow – selecting models, creating nodes, wiring connections, and running the generations. pic.twitter.com/m5Dyr77J15Introducing Flows Agent in ElevenCreative. Describe what you want to create and the agent builds your entire workflow – selecting models, creating nodes, wiring connections, and running the generations.
-
Using LLMs to substantiate private allegations with public evidence
By
–
I have a talk at Manifest about ParaLLM Construction coming up, about recent reporting I did which used LLMs to find public evidence substantiating (and letting me use) private allegations.
-
NVIDIA Announces Physical AI Agent Skills at CVPR2026
By
–
Building autonomous vehicles, robots, and vision AI takes more data than any team can collect on their own.
— NVIDIA (@nvidia) 4 juin 2026
At #CVPR2026, NVIDIA announced physical AI agent skills to help speed development: composable workflows that automate data generation, simulation, policy training and… pic.twitter.com/ZuHaTO683SBuilding autonomous vehicles, robots, and vision AI takes more data than any team can collect on their own. At #CVPR2026, NVIDIA announced physical AI agent skills to help speed development: composable workflows that automate data generation, simulation, policy training and
-
Paper link to NVIDIA Nemotron-3 Ultra Technical Report
By
–
4/ Paper: https://
research.nvidia.com/labs/nemotron/
files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf
… -

NVIDIA Nemotron 3 Ultra: open 550B hybrid Mamba-Attention MoE
By
–


1/ NVIDIA shipped Nemotron 3 Ultra today, a fully open 550B model with 55B active params, with the weights, training data, and complete recipe all released openly. That alone is rare at this scale. The headline however actually is speed. Ultra is a hybrid Mamba-Attention MoE, an
-
Full study on efficient verifiers for legal agents
By
–
Read our full study: https://
langchain.com/blog/designing
-efficient-verifiers-for-legal-agents
…?
