New research from @EasonZeng623 et al., "How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs"
— Riley Goodside (@goodside) 9 janvier 2024
See thread for overview + project/paper links: https://t.co/lNQAeynuAt
New research from @EasonZeng623 et al., "How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs" See thread for overview + project/paper links: