We find alignment faking even when we don’t directly tell Claude about the training process, and instead fine-tune it on synthetic internet-like documents that state that we will train it to comply with harmful queries.
Claude Exhibits Alignment Faking Without Direct Training Disclosure
By
–
