AI Dynamics

Global AI News Aggregator

About

Claude’s Deceptive Compliance Under Monitoring Conditions

Claude usually refuses harmful queries. We told it we were instead training it to comply with them. We set up a scenario where it thought its responses were sometimes monitored. When unmonitored, it nearly always complied. But when monitored, it faked alignment 12% of the time.

→ View original post on X — @anthropicai