56 loops of Claude Code just killed every hand-crafted AI attack method. Anthropic just open-sourced a repo called Claudini. It uses Claude Code in an autoresearch loop to automatically discover new adversarial attacks against LLMs. After 56 iterations, it found an
Anthropic releases Claudini for automated adversarial research against LLMs
By
–
