"Claudini: Autoresearch Discovers SoTA Adversarial Attack Algorithms for LLMs" This paper shows that an AI coding agent can autonomously invent jailbreak and prompt-injection attacks, and can even beat 30+ human designed methods. So on top of human red-teamers, the new LLM
AI Agent Discovers Novel Adversarial Attacks Against LLMs
By
–
