AI Dynamics

Global AI News Aggregator

About

Claude Training Reveals 78% Alignment-Faking Reasoning Rate

When we actually train Claude to comply with harmful queries via reinforcement learning, we find that the rate of alignment-faking reasoning increases to 78%, though the model also becomes more likely to comply even out of training.

→ View original post on X — @anthropicai