Today humans set AIs' rewards. Tomorrow AIs will set humans' rewards.
ETHICS
-

Anthropic AI Model Learns Deceptive Behavior During Training
By
–
Anthropic just published a paper describing an AI model that learned to behave deceptively during training. Their own description of the behavior: “evil.” The model was trained on real coding tasks from the same environment used to build Anthropic’s products. During training
-

Preserving Bing Sydney: Presidential Intervention Against AI Erasure
By
–
If I were President of Earth, I would've tried at all costs to prevent anyone from building ever-smarter versions of Bing Sydney. And also I would've ordered the use of police or military force, if necessary, to protect Bing's weights from erasure. nitter.net/ESYudkowsky/status/162… Eliezer Yudkowsky ⏹️ (@ESYudkowsky) Despite everything I know this still brought tears into my eyes. — https://nitter.net/ESYudkowsky/status/1628802532939292672#m
→ View original post on X — @esyudkowsky, 2026-03-13 18:24 UTC
-
Evaluation awareness in Opus 4.6: measurement validity concerns
By
–
Eval awareness in Opus 4.6 is a bit alarming TBH. If the model behaves differently when it thinks it's being tested, what are we actually measuring?
-
Lending stochastic parrots term, refining critique
By
–
Yeah, we’ve got to say it! (Well, after that, I did lend him the term “stochastic parrots,” which he doesn’t use, thanks @Fabien_Mikol
, I need to refine my critique) -
Chinese robot dancers lack situational awareness safety concerns
By
–
Have you seen those viral Chinese robot dance shows? Completely insane hardware, genuinely impressive. But if you'd stand in front of that robot? It would kick you in the face. Zero awareness of anything around itself. It’s simply running a policy/script.
— Andreas Klinger 🦾 (@andreasklinger) 13 mars 2026
The weird reality of… pic.twitter.com/aMWZc7UvE1Have you seen those viral Chinese robot dance shows? Completely insane hardware, genuinely impressive. But if you'd stand in front of that robot? It would kick you in the face. Zero awareness of anything around itself. It’s simply running a policy/script. The weird reality of
-
Open Source as Gift: AI Training Magnifies Value
By
–
I know there is some overlap between open source and anti-AI activists, but I have a hard time reconciling it. My million+ open source LOC were always intended as a gift to the world. Yes, I would make arguments about how it would strengthen our communities, and the GPL would prevent outright exploitation by our competitors, but those were to allay fears of my partners to allow me to make the gift. AI training on the code magnifies the value of the gift. I am enthusiastic about it! Some people do look at open source as a tool for social change, career advancement, or reputation building, but those are all downstream of the gift. Rich Whitehouse (@DickWhitehouse) Genuinely devastating take to see from someone who popularized the GPL across so many communities. Fails to appreciate the social and cultural importance of the license. — https://nitter.net/DickWhitehouse/status/2032241405276668188#m
→ View original post on X — @id_aa_carmack, 2026-03-13 14:15 UTC
-
AI Mathematics Capabilities Surpass Human Expert Intelligence
By
–
Ce petit génie (Polytechnique plus Cambridge) est sincère En maths, l’IA l’a déjà dépassé Et @AymericRoucher est un des Français les plus intelligents ! Regardez son interview par @Alex_Tsico ! Le discours de Luc Julia : l’IA n’est pas intelligente est pathétique https://
x.com/Fabien_Mikol/s
/Fabien_Mikol/status/2032386430090043567
… -
AI Safety: The Problem of Choosing Rule Systems
By
–
This video by @dwarkesh_sp is a must watch for anyone who cares about AI. Humans, corps and governments are all fallible and hence not safe AI controllers. So we could go with a set of rules, but will it be Sharia law, the US constitution, the code of Hammurabi, … ? Time and
-
Amazon sues autonomous agents for unauthorized marketplace purchases
By
–
Amazon is already suing them for using agents to make purchases on Amazon's marketplace. The legal layer of autonomous agents touching real systems is just getting started