Shared my first trace from @NanoClaw_AI to @huggingface yesterday. Very cool! By default, all agents should store their traces on HF (in private) so that you can keep a history of them, analyze them,… & share them and post-train better models, harnesses and more. Excited
CODE
-

AEP-001 GoalOS Proof-of-Evolution Constitution Standard
By
–
New standard for the agent era: AEP-001 — GoalOS Proof-of-Evolution Constitution Commit → Execute → Prove → Evolve. No proof, no evolution.
No eval, no propagation.
No rollback, no release. This is Proof-Carrying Intelligence. https://
montrealai.github.io/proof-gradient
/standards/AEP-001/
… #GoalOS -

AI21 Labs surpasses Claude Code in efficiency and performance with Test Agent.
By
–
4/5 Still came in ~$0.30 under Claude Code’s spend at a similar score. So we added a lightweight Test Agent that writes repo tests and filters failing patches, pushing our final result to 60.9% – surpassing Claude Code (60.9% vs 56.2%) at the same cost.
-

AI21 Labs: ReAct Agent Performance with Enrichment and Scaling Strategies
By
–
2/5 Started with a baseline: classic ReAct agent (GPT-5.2), single Docker-terminal tool. Baselines on the slice: vanilla 53.8%, enrich-only 55.6%, scale-only (n=5 + LLM judge) 55.4%, enrich-then-scale 57.7%.
-
LangSmith Engine reviews traces, learns from usage, updates Context Hub
By
–
You can use LangSmith Engine to review your agent traces to find bugs and areas for improvement across agent prompts + code. Between runs, the agent can review conversations, learn from real usage, and update Context Hub files.
-
MiniMax M3: Open-weights frontier model challenges closed model dominance
By
–
THE ERA OF RELYING EXCLUSIVELY ON THE 3 MAJOR CLOSED MODELS IS OVER@MiniMax_AI's M3 is officially out 💥💥💥
— Charly Wargnier (@DataChaz) 4 juin 2026
It delivers the exact same capabilities you expect from a frontier model, combining massive leaps forward in a highly cost-efficient, open-weights package.
Here's why… pic.twitter.com/NDUppZzMlqTHE ERA OF RELYING EXCLUSIVELY ON THE 3 MAJOR CLOSED MODELS IS OVER @MiniMax_AI
's M3 is officially out It delivers the exact same capabilities you expect from a frontier model, combining massive leaps forward in a highly cost-efficient, open-weights package. Here's why -
Claude’s neurosymbolic code useful, but more AI work needed
By
–
now/years. claude code is neurosymbolic and pretty useful in its domain, but there’s lots more to be done (see my 2020 article Next Decade in AI).
-
Claude mysteriously installs claude.exe on Ubuntu VPS
By
–
No idea how but Claude somehow installed claude.exe on my Ubuntu VPS
-

Launching SynthTraces: generating synthetic coding agent traces
By
–
Today I'm launching a new project called SynthTraces It is a minimal codebase to generate synthetic coding agent session traces using Pi (from @badlogicgames
) I wanted a large number of coding-agent traces, so I built a tiny harness where two models talk to each other: – an -
NVIDIA Nemotron 3 Ultra: Open Model for Agentic Tasks
By
–
Introducing NVIDIA Nemotron 3 Ultra.
— NVIDIA (@nvidia) 4 juin 2026
A frontier smart open model built for long-running agents that need to plan, reason, use tools and keep working across complex coding, research and enterprise workflows.
Up to 5x faster inference and up to 30% lower cost for agentic tasks.… pic.twitter.com/AcHTauUzjmIntroducing NVIDIA Nemotron 3 Ultra. A frontier smart open model built for long-running agents that need to plan, reason, use tools and keep working across complex coding, research and enterprise workflows. Up to 5x faster inference and up to 30% lower cost for agentic tasks.
