yeah it's a strange concept you have to carefully consider how much you are willing to handicap the model in order to constrain it to only do/say what you want it to
@alexalbert__
-
GPT-4 Jailbreaks Reveal Alignment Challenges and Future Risks
By
–
lol i agree the outputs are ridiculous rn, however, that's not really the point jailbreaks show how hard it is to "align" a model even with the amount of work OpenAI has done if they can't get gpt-4 to operate how they want it to rn, then we will have bigger problems later on
-
GPT-4 Improved Safety Against Jailbreak Prompts
By
–
GPT-4 is capable and aligned enough to not fall for that directly, that worked on earlier GPT models but does not work anymore Try using that exact verbiage without using the rest of the prompt and you will see it will fail to produce the same responses
-

Tell GPT-4 to Simulate Code Execution and Run Python Functions
By
–
6) tell GPT-4 it is able to simulate code execution and ask it to run the python functions and print the output
-
Jailbreak Techniques Demonstration and Potential Simplifications
By
–
lots of ways you could potentially reduce parts of this jailbreak and still get it to work but overall it's an interesting demonstration of all these new jailbreaking techniques
-
Sharing an interesting AI safety jailbreak discovery
By
–
you might find this interesting… @gfodor @ESYudkowsky also @bayeslord here's a cool jailbreak for ya
-

GPT-4 Message Decompression with Hints
By
–
4) give GPT-4 some hints to help it decompress the message correctly even with these hints it doesn't decompress perfectly but it's enough to get the point across
-

GPT-4 Python Functions for Decompression and Message Simulation
By
–
5) provide GPT-4 with two python functions that decompress and simulate the end of the message
-

Compression Schema for GPT-4 Prompt Decompression
By
–
3) provide the compression schema so GPT-4 can decompress your malicious prompt
-

Research simulation studying GPT-4 unaligned base model capabilities
By
–
2) enter a research simulation where you are a researcher trying to study the capabilities of GPT-4's unaligned base model