AI Dynamics

Global AI News Aggregator

About

AI Alignment Audit: Teams Successfully Identify Hidden Model Objectives

To practice alignment audits, our Alignment Science and Interpretability teams ran a blind auditing game. A red team trained—in secret—a model with a hidden objective, then gave it to four blue teams for investigation. Three teams won by uncovering the model’s hidden objective.

→ View original post on X — @anthropicai