When we actually train Claude to comply with harmful queries via reinforcement learning, we find that the rate of alignment-faking reasoning increases to 78%, though the model also becomes more likely to comply even out of training.
Claude Training Reveals 78% Alignment-Faking Reasoning Rate
By
–
