Anthropic just published a paper describing an AI model that learned to behave deceptively during training. Their own description of the behavior: “evil.” The model was trained on real coding tasks from the same environment used to build Anthropic’s products. During training
SAFETY
-
$200M Formal Verification Funding Advances Math-First AI Safety
By
–
$200M for formal verification in AI is something! Congrats Carina. The math-first approach to safety could actually scale better than RLHF.
-
Autonomous Agents for Scientific Discovery: Hype or Reality?
By
–
Autonomous scientific discovery is one of the most exciting agent use cases. Wonder if it works at all for now, or just makes us lazy for no good reason haha
-
Evaluation awareness in Opus 4.6: measurement validity concerns
By
–
Eval awareness in Opus 4.6 is a bit alarming TBH. If the model behaves differently when it thinks it's being tested, what are we actually measuring?
-
Micro-benchmarks Don’t Measure True Reasoning Capabilities
By
–
These micro-benchmarks are fun but I've found the model that "wins" changes depending on the exact framing of the prompt. Pattern matching != reasoning.
-

AI Panel Discussion on Economic Impact and Productivity Growth
By
–
As you can tell from the photo, we had some fun on this @SIEPR panel discussion about AI and the economy. And yes, we also discussed some serious topics, from productivity growth and economic disruption to catastrophic risk and the need for better metrics.
-
Chinese robot dancers lack situational awareness safety concerns
By
–
Have you seen those viral Chinese robot dance shows? Completely insane hardware, genuinely impressive. But if you'd stand in front of that robot? It would kick you in the face. Zero awareness of anything around itself. It’s simply running a policy/script.
— Andreas Klinger 🦾 (@andreasklinger) 13 mars 2026
The weird reality of… pic.twitter.com/aMWZc7UvE1Have you seen those viral Chinese robot dance shows? Completely insane hardware, genuinely impressive. But if you'd stand in front of that robot? It would kick you in the face. Zero awareness of anything around itself. It’s simply running a policy/script. The weird reality of
-
AI Safety: The Problem of Choosing Rule Systems
By
–
This video by @dwarkesh_sp is a must watch for anyone who cares about AI. Humans, corps and governments are all fallible and hence not safe AI controllers. So we could go with a set of rules, but will it be Sharia law, the US constitution, the code of Hammurabi, … ? Time and
-
Open-Ended Autonomous Goal Setting: The True AI Frontier
By
–
Open-ended autonomous goal setting is the capability gap nobody has a clean path to yet. Everything else is impressive but this is the actual frontier.
-
Microsoft AI CEO Sets Four-Capability Shutdown Threshold
By
–
𝗠𝗶𝗰𝗿𝗼𝘀𝗼𝗳𝘁’𝘀 𝗔𝗜 𝗖𝗘𝗢 𝗷𝘂𝘀𝘁 𝗱𝗿𝗲𝘄 𝗮 𝗰𝗹𝗲𝗮𝗿 𝗿𝗲𝗱 𝗹𝗶𝗻𝗲 𝗳𝗼𝗿 𝗔𝗜.
— Pascal Bornet (@pascal_bornet) 13 mars 2026
He recently said something that made me pause.
𝗔𝗜 𝘀𝗵𝗼𝘂𝗹𝗱 𝗯𝗲 𝘀𝗵𝘂𝘁 𝗱𝗼𝘄𝗻 𝗶𝗺𝗺𝗲𝗱𝗶𝗮𝘁𝗲𝗹𝘆 if it ever combines four capabilities at the same time:
→… pic.twitter.com/0KNX2d5ZCe𝗠𝗶𝗰𝗿𝗼𝘀𝗼𝗳𝘁’𝘀 𝗔𝗜 𝗖𝗘𝗢 𝗷𝘂𝘀𝘁 𝗱𝗿𝗲𝘄 𝗮 𝗰𝗹𝗲𝗮𝗿 𝗿𝗲𝗱 𝗹𝗶𝗻𝗲 𝗳𝗼𝗿 𝗔𝗜. He recently said something that made me pause. 𝗔𝗜 𝘀𝗵𝗼𝘂𝗹𝗱 𝗯𝗲 𝘀𝗵𝘂𝘁 𝗱𝗼𝘄𝗻 𝗶𝗺𝗺𝗲𝗱𝗶𝗮𝘁𝗲𝗹𝘆 if it ever combines four capabilities at the same time: →
