Detecting problems in agent traces in production is difficult. It must be done at low cost (due to volume) but also with accuracy (otherwise too much noise). We post-trained our own model for this. SOTA accuracy, at ~10-100x lower cost.
RESEARCH
-

Models weak on vision cause error accumulation in visual steps
By
–
Very clever. And matches what I would expect: models are weak on vision relative to everything else, so visual steps are where errors accumulate most in workflows.
-
Model distillation: the most powerful are American, interesting examples
By
–
The reality is that everyone distills from everyone, usually from the most powerful, and the most powerful right now are American models. Here are some interesting examples.
-
Stanford HAI director Fei-Fei Li featured on FastCompany cover explaining world models
By
–
HAI Founding Director @drfeifei is featured on @FastCompany
's cover, explaining "world models" – AI that understand physical space and real-world dynamics. Rooted in human-centered philosophy, she explains what makes it different and what's at stake: -
The biggest obstacles in building scientific agents according to nlarusstone
By
–
In the latest Max Agency, @hwchase17 asked @benchling Head of AI @nlarusstone about the biggest blockers in building agents for scientific work. pic.twitter.com/XKS6Nnj5mv
— LangChain (@LangChain) 15 juin 2026In the latest Max Agency, @hwchase17 asked @benchling's Head of AI @nlarusstone what were the biggest obstacles in building agents for scientific work.
-
Train LLM from scratch and AI Engineering book resources
By
–
Train LLM from scratch: https://
github.com/FareedKhan-dev
/train-llm-from-scratch
…
—
AI Engineering book: https://
dailydoseofds.github.io/ai-engg-book/ -

Build a GPT-style transformer from scratch without high-level libraries
By
–
Train your own LLM from scratch. This repo builds a GPT-style transformer from the ground up, without using any high-level libraries. You see exactly how attention, multi-head attention, the feed-forward block, embeddings, residuals, and layer norm fit together. And it doesn't
-

AI solves 7/10 hard math problems but still criticized
By
–

Weird headline – I am not sure solving 7 out of 10 novel very hard problems meant AI "did not live up to the task," when 15 months ago LLMs couldn't do math. But the actual study is interesting and illuminates flaws & successes of AIs in math. https://
1stproof.org/assets/docs/re
port.pdf
… -
Concern about a possible leak of Claude and Codex projects
By
–
Read carefully. What worries me most is a possible leak one day of the projects worked on in the Claude and Codex environments.
-
Open models enable AI public good contributions from non-frontier nations.
By
–
Current open models are now good enough to pull off some of these projects if scaffolded properly while others (co-scientist) benefit from the AI frontier. Because of that, this is an area where nations without frontier labs could contribute to the impact of AI for public good.
