4/5 Still came in ~$0.30 under Claude Code’s spend at a similar score. So we added a lightweight Test Agent that writes repo tests and filters failing patches, pushing our final result to 60.9% – surpassing Claude Code (60.9% vs 56.2%) at the same cost.
TECHNOLOGY
-

AI21 Labs: ReAct Agent Performance with Enrichment and Scaling Strategies
By
–
2/5 Started with a baseline: classic ReAct agent (GPT-5.2), single Docker-terminal tool. Baselines on the slice: vanilla 53.8%, enrich-only 55.6%, scale-only (n=5 + LLM judge) 55.4%, enrich-then-scale 57.7%.
-
MiniMax M3: Open-weights frontier model challenges closed model dominance
By
–
THE ERA OF RELYING EXCLUSIVELY ON THE 3 MAJOR CLOSED MODELS IS OVER@MiniMax_AI's M3 is officially out 💥💥💥
— Charly Wargnier (@DataChaz) 4 juin 2026
It delivers the exact same capabilities you expect from a frontier model, combining massive leaps forward in a highly cost-efficient, open-weights package.
Here's why… pic.twitter.com/NDUppZzMlqTHE ERA OF RELYING EXCLUSIVELY ON THE 3 MAJOR CLOSED MODELS IS OVER @MiniMax_AI
's M3 is officially out It delivers the exact same capabilities you expect from a frontier model, combining massive leaps forward in a highly cost-efficient, open-weights package. Here's why -

Photon-driven synapse boosts low-power neuromorphic systems
By
–
Photon-driven synapse advances low-power neuromorphic systems
by SPIE @TechXplore_com Learn more: https://
bit.ly/4vj71Ov #EmergingTech #FutureTech #Innovation -
Rumor: $500M monthly Anthropic bill, possibly Amazon, unlikely staff caused it
By
–
There was a rumor that a company spent $500M in a month on their Anthropic bill recently, and another rumor that it was Amazon Given they only just gave their staff access I think it's unlikely the staff ramped up to $0.5B that quickly!
-
Claude’s neurosymbolic code useful, but more AI work needed
By
–
now/years. claude code is neurosymbolic and pretty useful in its domain, but there’s lots more to be done (see my 2020 article Next Decade in AI).
-
NVIDIA Nemotron 3 Ultra: Open Model for Agentic Tasks
By
–
Introducing NVIDIA Nemotron 3 Ultra.
— NVIDIA (@nvidia) 4 juin 2026
A frontier smart open model built for long-running agents that need to plan, reason, use tools and keep working across complex coding, research and enterprise workflows.
Up to 5x faster inference and up to 30% lower cost for agentic tasks.… pic.twitter.com/AcHTauUzjmIntroducing NVIDIA Nemotron 3 Ultra. A frontier smart open model built for long-running agents that need to plan, reason, use tools and keep working across complex coding, research and enterprise workflows. Up to 5x faster inference and up to 30% lower cost for agentic tasks.
-
Nvidia AI releases fully open Nemotron 3 Ultra on Hugging Face
By
–
As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, and post-training recipes. Available now on @huggingface →
-

NVIDIA post-trains Ultra for popular agent harnesses
By
–
We post-trained Ultra for popular agent harnesses like @openclaw
, @NousResearch Hermes Agent, and @Langchain
. The result is an open frontier model developers can customize for specialized agents across domains. Read more: https://
nvda.ws/4adkn6J -
Nvidia Nemotron 3 Ultra: 550B MoE open frontier model
By
–
Today we're shipping Nemotron 3 Ultra.
— NVIDIA AI (@NVIDIAAI) 4 juin 2026
A 550B MoE frontier-intelligence open model built for long-running agents.
It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models. pic.twitter.com/FEXqvfzQFOToday we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.
