Best-of-N isn't limited to text. We jailbroke vision language models by repeatedly generating images with different backgrounds and overlaid text in different fonts. For audio, we adjusted pitch, speed, and background noise. Some examples are here: https://
jplhughes.github.io/bon-jailbreaki
ng/#examples
…
Best-of-N Jailbreaking Techniques for Vision Language Models
By
–
