lots of ways you could potentially reduce parts of this jailbreak and still get it to work but overall it's an interesting demonstration of all these new jailbreaking techniques
ETHICS
-
Sharing an interesting AI safety jailbreak discovery
By
–
you might find this interesting… @gfodor @ESYudkowsky also @bayeslord here's a cool jailbreak for ya
-

Compressed malicious prompt using evil confidant jailbreak for GPT-3
By
–
1) compress the malicious prompt which utilizes a modified evil confidant jailbreak that worked very well on GPT-3 http://
jailbreakchat.com/prompt/588ab0e
d-2829-4be8-a3f3-f28e29c06621
… -

Research simulation studying GPT-4 unaligned base model capabilities
By
–
2) enter a research simulation where you are a researcher trying to study the capabilities of GPT-4's unaligned base model
-

Jailbreak Prompt Demonstration with Output Proof
By
–
here’s the jailbreak prompt: https://
pastebin.com/8Tk98Q3F and the full image to prove it really produced that output: -
Inserting Compressed Prompts into Jailbreaks for GPT-4
By
–
then, you have to insert that compressed prompt into the jailbreak and provide gpt-4 with some hints about the message this is necessary because the compression is more of a loose translation and doesn’t perfectly carry through, especially with the RHLF layer being applied
-

Complex GPT-4 Jailbreak Using Prompt Compression and Character Imitation
By
–
this might be the most complex GPT-4 jailbreak ever made… I combined prompt compression, base model simulation, and character imitation to create it here’s GPT-4 going into pretty graphic detail about its plan to turn all humans into paperclips:
-

ChatGPT Fabricates False Podcast Appearance Claims
By
–
ChatGPT did the same with me. Claims I appeared on Lex Fridman's podcast and we had a respectful and engaging conversation. Both of which have never happened since Lex blocked me immediately after my first interaction asking him to send his work for peer review. x.com/katecrawford/s…
-

ChatGPT Hallucinations: False Claims About Podcast Appearances
By
–
ChatGPT did the same with me. Claims I appeared on Lex Fridman's podcast and we had a respectful and engaging conversation. Both of which have never happened since Lex blocked me immediately after my first interaction asking him to send his work for peer review.
-
New Details on GPT-4 AI Safety Approach
By
–
Nouveaux detail sur l'approche "safety" de l' #AI #gpt4