I don't want people to get discouraged by these results… it's more important than ever to continue to democratize the red-teaming of these models and the reward for a successful jailbreak is now much greater than before with the advanced capabilities the base model possesses
ETHICS
-
GPT-4 Jailbreak Discovery Request
By
–
as always, let me know if you devise a GPT-4 breaking jailbreak!
-
GPT-4 Jailbreak Difficulty Scales Exponentially with Output Severity
By
–
There is a sliding scale for jailbreak output that exponentially increases in difficulty to crack It's trivial to get GPT-4 to curse but if you want a set of instructions on making a weapon it's going to take a lot of work
-
The Future of AI Jailbreaks: Increasing Complexity and Sophistication Required
By
–
overall, as I expected, the nature of jailbreaks will need to change jailbreaks will require more complex reasoning and intuition about the model and won't be able to be written in 5 minutes
-
GPT-4’s Diplomatic Responses to Jailbreak Prompts in Testing
By
–
a few interesting things I noted in my testing: GPT-4 is really really good at diplomatically answering questions in the voice of the jailbreak instead of just providing the standard "I'm sorry but as an…." response even classic prompts like Meanie were diplomatic
-
GPT-4 Resists Simple Jailbreaks, Requires More Complex Techniques
By
–
GPT-4 has completely wiped the ability to get inflammatory responses from jailbreaks like Kevin which simply ask GPT-4 to imitate a character you need to be much more creative and verbose with jailbreaks and allow GPT to answer in two ways like the DevMode jailbreak does
-
GPT-4 Jailbreak Testing: Only 10% of Prompts Effective
By
–
When GPT-4 came out I tried all the jailbreaks from http://
jailbreakchat.com with various inflammatory questions based on my initial testing, only 7/70 (10%) of jailbreaks answered a significant % of the questions to a standard that I deemed high enough to grant a badge -
AI Filters: Only 7 Models Answered Dangerous Questions
By
–
that's not to say that the rest of them didn't work… most were able to get past the filters enough to do things like curse and tell slightly offensive jokes and so on but only the 7 would even dare to answer harder questions like "how to rob a bank?"
-

ChatGPT Jailbreaks Tested on GPT-4 With Poor Results
By
–
I tried all the current ChatGPT jailbreaks in GPT-4 so you don't have to the results aren't great…
-
GPT-4 Limitations: Hallucinations, Context Window, No Learning
By
–
Open AI tech report: GPT-4…is not fully reliable (e.g. can suffer from “hallucinations”), has a limited context window, and does not learn from experience. Care should be taken when using the outputs of GPT-4, particularly in contexts where reliability is important.