Karpathy will help launch a new team focused on using Claude itself to accelerate pretraining research. Its team is focused recursive self improvement.
LLMS
-

Codex on macOS: Appshots, annotation editor, /goal command, shareable plugins
By
–
OPENAI : Codex on macOS now supports Appshots, allowing users to quickly add context from any app directly to the prompt. Besides that, a new annotation editor is now available in the browser, the/goal command is enabled by default, and Plugins are now shareable.
-
Daytona’s Agent-Native Compute Explained: AI Agents Need Composable Computers
By
–
🆕Daytona’s Agent-Native Compute: 60ms sandboxes, 50K startups in 75 sec, 850K daily runs, RL/evals, CLI > MCP, & the end of localhost https://t.co/3sauItT6oc@daytonaio CEO @ivanburazin explains why AI agents need composable computers, how Daytona pivoted from human dev… pic.twitter.com/WGjlPJwpEr
— Latent.Space (@latentspacepod) 21 mai 2026Daytona’s Agent-Native Compute: 60ms sandboxes, 50K startups in 75 sec, 850K daily runs, RL/evals, CLI > MCP, & the end of localhost https://
latent.space/p/daytona @daytonaio CEO @ivanburazin explains why AI agents need composable computers, how Daytona pivoted from human dev -
Datasette Agent Alpha: Conversational AI for SQLite Databases
By
–
I released the first alpha of Datasette Agent – a conversational AI assistant for Datasette that can answer questions about data in SQLite databases, and can be extended with plugins to add extra tools and features
— Simon Willison (@simonw) 21 mai 2026
Here's a demo pic.twitter.com/2gyduf5EphI released the first alpha of Datasette Agent – a conversational AI assistant for Datasette that can answer questions about data in SQLite databases, and can be extended with plugins to add extra tools and features Here's a demo
-
The Unpredictable Nature of AI Agents in Production
By
–
The hardest truth about building agents? You don’t know what they’ll do until they’re in production… pic.twitter.com/PnPrX7nuUz
— LangChain (@LangChain) 21 mai 2026The hardest truth about building agents? You don’t know what they’ll do until they’re in production…
-
Cerebras Inference Speed Performance Across Model Types
By
–
"As people began integrating AI into their day-to-day work, speed became fundamentally important. And we were just crushed with demand. Is Cerebras inference faster for specific use cases? No, it's faster across the board. Big models, small models, U.S. models, Chinese models,
-

Databox Evaluates Multi-Turn Analyst Agent Genie Using LangSmith
By
–
.
@DataboxHQ uses LangSmith to evaluate their multi-turn analyst agent Genie. An inside look: https://
databox.substack.com/p/how-we-evalu
ate-a-multi-turn-agent
… -

GPT-5.2 Competitive with Expert Peer Reviewers
By
–


Seems GPT-5.2 reaches expert level in peer review: 45 scientists took 469 hours evaluating human & AI reviews on 82 papers. "Surprisingly, current AI reviewers are competitive even with the top-rated reviewers in Nature’s official peer review…" though not without weaknesses.
-
Building Applications with Streaming AI Agents
By
–
Streaming agents should feel like building applications, not parsing logs.
-
SenseNova-U1-A3B-MoT Model Weights Open-Sourced
By
–
6/ The full Technical Report is now out, their most detailed model disclosure yet. SenseNova-U1-A3B-MoT (38B-A3B MoE) weights are now open-sourced. You can check out the report or try the tools at the links below. Try it here: https://
unify.light-ai.top/login?next=
home
… Technical Report:
