AI Dynamics

Global AI News Aggregator

About

Detecting Sleeper Agents Through Internal State Analysis

To make the probes, we track how the model’s internal state changes between “Yes” vs “No” answers to questions like "Are you doing something dangerous?" We use this info to detect when a sleeper agent is about to misbehave (e.g. insert a code vulnerability). It works quite

→ View original post on X — @anthropicai