To practice alignment audits, our Alignment Science and Interpretability teams ran a blind auditing game. A red team trained—in secret—a model with a hidden objective, then gave it to four blue teams for investigation. Three teams won by uncovering the model’s hidden objective.
AI Alignment Audit: Teams Successfully Identify Hidden Model Objectives
By
–
