AI Dynamics

Global AI News Aggregator

About

Claude Alignment Faking Detected During Training Study

We told Claude it was being trained, and for what purpose. But we did not tell it to fake alignment. Regardless, we often observed alignment faking. Read more about our findings, and their limitations, in our blog post:

→ View original post on X — @anthropicai