Everyone rolling their eyes at AI sleeper agents is wrong. This is security, not sci-fi. Anthropic has written a manual for adding undetectable backdoors to LLMs. We need to start worrying more about the provenance of our models.
AI Sleeper Agents: Security Threat from Undetectable LLM Backdoors
By
–