AI Dynamics

Global AI News Aggregator

About

Anthropic Research: Probes Detect Backdoored Sleeper Agent Models

New Anthropic research: we find that probing, a simple interpretability technique, can detect when backdoored "sleeper agent" models are about to behave dangerously, after they pretend to be safe in training. Check out our first alignment blog post here: https://
anthropic.com/research/probe
s-catch-sleeper-agents

→ View original post on X — @anthropicai