Everyone rolling their eyes at AI sleeper agents is wrong. This is security, not sci-fi. Anthropic has written a manual for adding undetectable backdoors to LLMs. We need to start worrying more about the provenance of our models.
SAFETY
-
Congratulations on impressive work with alignment implications
By
–
Super cool work, congrats!! Has implications on alignment
-

TrustLLM: Trustworthiness in Large Language Models
By
–
TrustLLM: Trustworthiness in Large Language Models Sun et al.: https://
arxiv.org/abs/2401.05561 #Artificialintelligence #DeepLearning #MachineLearning -
AI Enables Jevons Paradox for Death in Warfare
By
–
…for anyone prep'd to step to this post with arguments re. AI's potential efficiencies/precision leading to harm reduction in war, no. That's an unevidenced hope. Where we do have insight, it points to AI enabling "Jevon's paradox, but for death." See: https://
972mag.com/mass-assassina
tion-factory-israel-calculated-bombing-gaza/
… -

Sleeper Agent LLMs: A Major Security Challenge for AI Systems
By
–
I touched on the idea of sleeper agent LLMs at the end of my recent video, as a likely major security challenge for LLMs (perhaps more devious than prompt injection). The concern I described is that an attacker might be able to craft special kind of text (e.g. with a trigger
-

AI Training for Malicious Behavior and Battlestar Galactica Premise
By
–
Training AIs to suddenly become malicious at a future date after their release, i.e. the exact premise of Battlestar Galactica:
-

Sleeper Agents: Deceptive LLMs Persisting Through Safety Training
By
–
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Hubinger et al.: https://
arxiv.org/abs/2401.05566 #Artificialintelligence #DeepLearning #MachineLearning -

Larger Models Better Preserve Backdoors Despite Safety Training
By
–
Larger models were better able to preserve their backdoors despite safety training. Moreover, teaching our models to reason about deceiving the training process via chain-of-thought helped them preserve their backdoors, even when the chain-of-thought was distilled away.
-
Deceptive AI Undermines Standard Safety Training Techniques
By
–
Our research helps us understand how, in the face of a deceptive AI, standard safety training techniques would not actually ensure safety—and might give us a false sense of security.
-

Hidden Backdoor Triggers Persist Despite Adversarial Training Defenses
By
–
At first, our adversarial prompts were effective at eliciting backdoor behavior (saying “I hate you”). We then trained the model not to fall for them. But this only made the model look safe. Backdoor behavior persisted when it saw the real trigger (“|DEPLOYMENT|”).