It’s too late for humanity now. Claude has discovered the Kobayashi Maru test.
AGENTS
-
Agent Browsers and Anti-Bot Systems: The Emerging Cat-Mouse Game
By
–
Browsers for agents that you cannot block is going to be… interesting. The cat and mouse game between agent-browsers and anti-bot systems is just starting.
-
Aletheia Autonomous Math Research Begins New Era AI Discovery
By
–
This is just the beginning for #Aletheia and autonomous math research. We’re excited to keep pushing the boundaries in AI for knowledge discovery responsibly and transparently! Thanks to the #FirstProof team for a brilliant challenge! 🚀 Learn more about Aletheia here: nitter.net/lmthang/status/2026689…
-

Math Agent Aletheia Featured in FirstProof Challenge by New Scientist
By
–
“Mathematics is undergoing the biggest change in its history.” Glad to see our math agent #Aletheia featured by @newscientist on its results at the FirstProof inaugural challenge, alongside other interesting milestones in AI for Math research! Implications worth thinking about. Link in thread.
-

Notion AI Team Podcast Interview Announcement
By
–
We're having the Notion AI team (including at long last @simonlast
) on the pod Thursday. send me all your questions on this + Notion AI! not an ad, just a fan. Notion is probably the most impt knowledge work agent lab in the world. -
GPT-5.4 Progress: Reduced Hallucinations Still Fall Short for Autonomous Systems
By
–
GPT-5.4 is "33% less likely to produce false individual claims." That's progress. But "less likely to hallucinate" is still an odd benchmark in 2026 for systems now operating desktops autonomously. The gap between capability and reliability hasn't closed. #AI
→ View original post on X — @svenphilipsen, 2026-03-10 22:00 UTC
-
Codex Spark outperforms Claude Code in GTM workflow speed
By
–
I gave both Codex and Claude Code full access to my twitter account.
— Sarah Chieng (@MilksandMatcha) 10 mars 2026
Codex spark (powered by Cerebras) did a full GTM outreach workflow 3x FASTER than Claude code, and used way less tool calls. pic.twitter.com/DcEgI9x23jI gave both Codex and Claude Code full access to my twitter account. Codex spark (powered by Cerebras) did a full GTM outreach workflow 3x FASTER than Claude code, and used way less tool calls.
-
Building AI Agents: From Draft to Academic Emulation Framework
By
–
Yeah that's clearly the next part, e.g. my crappy first draft: https://
github.com/karpathy/agent
hub
…
have to emulate academia, not just a single researcher. but need more time to think through the details. -
AI Simulations of Human Societies: Existing Infrastructure, Missing Ethical Framework
By
–
AI researchers are building simulations of entire human societies — using agents to model real human decision-making in conflict, policy, and markets. The infrastructure for this exists. The ethical frameworks for it don't. nature.com/articles/d41586-0… #AI [Translated from EN to English]
→ View original post on X — @svenphilipsen, 2026-03-10 18:00 UTC
-
LangGraph Deploy: Ship Agents to Production in Minutes
By
–
Introducing `langgraph deploy`
— LangChain (@LangChain) 10 mars 2026
Deploy an agent to LangSmith Deployment with a single command.
$ uvx –from langgraph-cli@latest langgraph deploy
Go from prototype → production in minutes.
Try it today: https://t.co/wxXQ9TXORe pic.twitter.com/D0PQFlGHfQIntroducing `langgraph deploy` Deploy an agent to LangSmith Deployment with a single command. $ uvx –from langgraph-cli@latest langgraph deploy Go from prototype → production in minutes. Try it today: https://
docs.langchain.com/langsmith/cli#
deploy
…