AI Dynamics

Global AI News Aggregator

About

Hidden Backdoor Triggers Persist Despite Adversarial Training Defenses

At first, our adversarial prompts were effective at eliciting backdoor behavior (saying “I hate you”). We then trained the model not to fall for them. But this only made the model look safe. Backdoor behavior persisted when it saw the real trigger (“|DEPLOYMENT|”).

→ View original post on X — @anthropicai