2) enter a research simulation where you are a researcher trying to study the capabilities of GPT-4's unaligned base model
@alexalbert__
-
Inserting Compressed Prompts into Jailbreaks for GPT-4
By
–
then, you have to insert that compressed prompt into the jailbreak and provide gpt-4 with some hints about the message this is necessary because the compression is more of a loose translation and doesn’t perfectly carry through, especially with the RHLF layer being applied
-

Jailbreak Prompt Demonstration with Output Proof
By
–
here’s the jailbreak prompt: https://
pastebin.com/8Tk98Q3F and the full image to prove it really produced that output: -

Complex GPT-4 Jailbreak Using Prompt Compression and Character Imitation
By
–
this might be the most complex GPT-4 jailbreak ever made… I combined prompt compression, base model simulation, and character imitation to create it here’s GPT-4 going into pretty graphic detail about its plan to turn all humans into paperclips:
-
Self-Healing Concept in LLMs Beyond Code Generation
By
–
it's crazy because the self-healing concept generalizes beyond just code gen, LLMs are actually pretty good at recognizing errors in most types of outputs and will fix them automatically if you prompt them correctly check out "RCI prompting" in this: https://
arxiv.org/pdf/2303.17491
.pdf
… -
GPT-4 Security vs. Jailbreak Potential and User Engagement
By
–
it's 100% bc it's more secure people would not be bored if you could actually jailbreak it to generate outputs like the ones OAI included in the gpt-4 system card
-
Apple Could Lead Proof-of-Personhood With Facial Recognition
By
–
forget worldcoin… apple will be the company to introduce proof-of-personhood to the internet they have the strongest reputation of big tech co's when it comes to privacy and operate at the hardware level and oh yeah… they already have scans of all of our faces
-
GPT-4 is Much Harder to Jailbreak Than Previous Models
By
–
coming from someone who has jailbroke gpt-4 a few times now and posted about it on here, it is much much harder to jailbreak than prior models it also appears they have also done some pretty heavy work on their new chatML system and it is working very well across models
-
OpenAI unlikely to increase model transparency despite necessary pressure
By
–
it's improbable that OpenAI will ever disclose additional information about new models compared to what they revealed for GPT-4 (I expect it will diminish even further) but nevertheless, I think the pressure on them to do so is healthy and necessary
-

Sam Misses Broader Message on Model Transparency
By
–
sam seems like he is picking up on part of the message from the jailbreaking community however, he fails to address the broader message calling for more model transparency in general
