I mentioned this project in this week's AI Lab newsletter. It suggests a kind of flip side to the idea of self-improvement in AI models. As models become more powerful, it shows they can automate the process of figuring out how to jailbreak themselves and other models remarkably
AI Models Automate Self-Jailbreaking Process Advancement
By
–
