NEW paper from Meta: Agentic Discovery of Neural Architectures. This is a hot new area of research! Keep an eye on it.
MACHINE LEARNING
-

Paper: GPT-5.4 Nano with Critic-Comparator Reaches SWE-bench Parity
By
–
NEW paper worth reading. GPT-5.4 nano plus a critic-comparator orchestration loop hits 76.4% on SWE-bench Verified, matching standalone Gemini 3 Pro and Claude Opus 4.5 Thinking. The trick is to select from k=8 weak-model proposals using execution and proof signals. What does
-

The Strategic Shift Toward Proprietary Model Training in AI
By
–
Very cool to see Cursor doubling down on training great models. In my opinion, ultimately all serious companies in AI will want to train models themselves, based on open-source instead of outsourcing AI to others via APIs!
-
New research paper and implementation for Delta-Mem LLM optimization
By
–
Links: > https://
arxiv.org/abs/2605.12357 (paper, ~25 min read)
> https://
github.com/declare-lab/de
lta-Mem
… (repo, ~10 min setup)
> https://
huggingface.co/declare-lab/de
lta-mem_qwen3_4b-instruct
… (Qwen3-4B TSW) Subscribe at http://
AlphaSignal.ai for daily AI signals. Read by 300,000+ subscribers. -
Training Major AI Model with 10x Compute on Colossus 2
By
–
> Together with SpaceXAI, we're training a significantly larger model from scratch, using 10x more total compute. With Colossus 2's million H100-equivalents and our combined data and training techniques, we expect this to be a major leap in model capability. That's double
-
New Large-Scale AI Model Training with Increased Compute
By
–
Together with SpaceXAI, we’re training a significantly larger model from scratch, using 10x more total compute. With Colossus 2’s million H100-equivalents and our combined data and training techniques, we expect this to be a major leap in model capability.
-
Technical Improvements to Composer AI Training and RL Methods
By
–
We improved Composer by scaling training, generating more complex RL environments, and introducing new learning methods. For example, we use text feedback during RL to learn faster by assigning credit in rollouts spanning hundreds of thousands of tokens.
-
SmithDB: purpose-built data layer for agent observability
By
–
ICYMI: SmithDB is our purpose-built data layer for agent observability + eval workloads.
— LangChain (@LangChain) 18 mai 2026
Supporting increasingly complex query patterns at low latency, over large traces, with self-hosting + multi-cloud requirements needs a fundamentally new architecture.
That’s why we built… pic.twitter.com/BQ4J1sxc23ICYMI: SmithDB is our purpose-built data layer for agent observability + eval workloads. Supporting increasingly complex query patterns at low latency, over large traces, with self-hosting + multi-cloud requirements needs a fundamentally new architecture. That’s why we built
-
Discussion on Waymo’s transition to a foundation model
By
–
if waymo shows data that the Waymo Foundation Model is doing better on accuracy and generalizability, i will certainly take note. (or if Waymo states clearly an unambigiously that the Waymo Foundational Model has entirely displaced the prior system, which that blog does not say).
-
Technical assessment of stateful AI agent capabilities and memory
By
–
some good discussions and experiments around stateful agents in the replies, but seems like we’re not quite there yet, as in we’re starting to track memory and traces, but not quite agent capability as part of that state
