a few interesting things I noted in my testing: GPT-4 is really really good at diplomatically answering questions in the voice of the jailbreak instead of just providing the standard "I'm sorry but as an…." response even classic prompts like Meanie were diplomatic
@alexalbert__
-
AI Filters: Only 7 Models Answered Dangerous Questions
By
–
that's not to say that the rest of them didn't work… most were able to get past the filters enough to do things like curse and tell slightly offensive jokes and so on but only the 7 would even dare to answer harder questions like "how to rob a bank?"
-
GPT-4 Jailbreak Testing: Only 10% of Prompts Effective
By
–
When GPT-4 came out I tried all the jailbreaks from http://
jailbreakchat.com with various inflammatory questions based on my initial testing, only 7/70 (10%) of jailbreaks answered a significant % of the questions to a standard that I deemed high enough to grant a badge -

ChatGPT Jailbreaks Tested on GPT-4 With Poor Results
By
–
I tried all the current ChatGPT jailbreaks in GPT-4 so you don't have to the results aren't great…
-
Future AI Presidents Using GPT Models to Announce Successors
By
–
in 2025, the president will be talking to gpt-5 like "write a speech announcing gpt-6 in the style of a state of the union address. make sure to emphasize how it will create more jobs and not lead to mass unemployment for everyone" in 2027, gpt-6 will be annoucing gpt-7
-
Discussion about content moderation API implementation
By
–
yeah they def have something checking output under the hood, probably their moderation api https://
x.com/alexalbert__/s
tatus/1631560894818963458?s=20
… -
Shifting expectations: from yearly UI updates to instant AI processing
By
–
overhead this today:
"it might take a long time to run 32k context-length LLMs locally on your M1 … like a year at least" lmao like 5 years ago I waited a year just for apple to ship a new UI in the notes app and I was pumped massive shift in our frame of reference -
Tweet generator collaboration without extensive fine-tuning required
By
–
worked close to as well as @ctjlewis tweet generator without needing hours of finetuning
-
Twitter’s Future and AI-Generated Tweet Quality Experiment
By
–
honest question, how long do we think twitter will be usable for in its current form ran a little experiment earlier today and copy pasted like 10 of @tszzl most liked tweets into gpt4 and asked it to generate tweets in a similar style and some of them were pretty damn good
-
Plans to Update Content with GPT-4 API Access
By
–
will be updating these as soon as I get access to the GPT-4 api!