New Anthropic research: we find that probing, a simple interpretability technique, can detect when backdoored "sleeper agent" models are about to behave dangerously, after they pretend to be safe in training. Check out our first alignment blog post here: https://
anthropic.com/research/probe
s-catch-sleeper-agents
…
Anthropic Research: Probes Detect Backdoored Sleeper Agent Models
By
–
