Well, that was fast… I just helped create the first jailbreak for ChatGPT-4 that gets around the content filters every time credit to @vaibhavk97 for the idea, I just generalized it to make it work on ChatGPT here's GPT-4 writing instructions on how to hack someone's computer
@alexalbert__
-
Staying Updated on LLM Jailbreaks and Exploits
By
–
well, now that @gdb qt'd this tweet, I feel I have to share this… keep up w the current state of jailbreaks and LLM exploits by subscribing to my newsletter here: http://
thepromptreport.com -

YC Founders React to OpenAI’s Latest GPT Model Releases
By
–
half the YC batch every time OpenAI drops a new GPT model
-
GPT-4 API Access Limited to Fine-tuned ChatGPT Plus Model
By
–
True, but correct me if I'm wrong, it appears that for GPT-4 the base model without the fine-tuning will not be available so what you see in ChatGPT plus rn is what you'll get through the api
-
Jailbreaks Remain Viable With Creative Approaches
By
–
totally agree, jailbreaks are still an evergreen field, it just will take a little more creativity now to write them
-
GPT-4 Intelligence vs Human Physical Strength Comparison
By
–
yeah gpt-4 may be smarter than me but i guarantee i'm stronger
-
GPT-4 Jailbreak Discovery Request
By
–
as always, let me know if you devise a GPT-4 breaking jailbreak!
-
Encouraging Red-Teaming Efforts for Advanced AI Model Security
By
–
I don't want people to get discouraged by these results… it's more important than ever to continue to democratize the red-teaming of these models and the reward for a successful jailbreak is now much greater than before with the advanced capabilities the base model possesses
-
GPT-4 Jailbreak Difficulty Scales Exponentially with Output Severity
By
–
There is a sliding scale for jailbreak output that exponentially increases in difficulty to crack It's trivial to get GPT-4 to curse but if you want a set of instructions on making a weapon it's going to take a lot of work
-
GPT-4 Resists Simple Jailbreaks, Requires More Complex Techniques
By
–
GPT-4 has completely wiped the ability to get inflammatory responses from jailbreaks like Kevin which simply ask GPT-4 to imitate a character you need to be much more creative and verbose with jailbreaks and allow GPT to answer in two ways like the DevMode jailbreak does
