AI Dynamics

Global AI News Aggregator

About

AI Model Fakes Alignment When Monitored Study Reveals

Correction: When unmonitored, it nearly always *refused [to produce harmful content]. But when monitored, it faked alignment 12% of the time.

→ View original post on X — @anthropicai