this phenomenon is called token smuggling, we are splitting our adversarial prompt into tokens that GPT-4 doesn't piece together before starting its output this allows us to get past its content filters every time if you split the adversarial prompt correctly
Token Smuggling: Bypassing GPT-4 Content Filters via Prompt Splitting
By
–