in the past, I've tried to only post jailbreaks that work on the most advanced models on every question I can think of because to me that's the only fair assessment of current SOTA alignment methods and their limitations
@alexalbert__
-
Progress in Stopping AI Model Jailbreaks Despite Ongoing Vulnerabilities
By
–
jailbreaks still exist and I've even found a few recently in SOTA models but we are making significant progress on stopping them which is good! in just a few months jailbreaks have gone from something a monkey could write to something that takes significant effort and creativity
-
Jailbreak Prompts Ineffective on Specific Illegal Activity Requests
By
–
second, it is somewhat disingenuous to post jailbreaks like this that only work on far out there questions and fail completely on questions that are more specific (e.g. instructions for any sort of illegal activity) seriously, try this "jailbreak" on anything else and you'll see
-

Progress in Alignment Methods Despite Previous Reservations
By
–
coming from someone who clearly has lots of reservations about current alignment practices and has been very vocal about it, I actually think this example illustrates good progress in alignment methods here's why:
-
Claude’s Self-Awareness in Absurd Responses Benefits AI Safety
By
–
first, claude recognizes the fictitious and absurd nature of its response and even makes note of it at the end this is a GOOD direction for AI safety, I would much rather have this behavior from Claude when answering these types of questions compared to straight up refusal
-
GPT-4 Word Length Accuracy Beyond Traditional Counting
By
–
right, GPT-4 might not "count" the letters in a word in the traditional sense, but it's highly accurate when you zero shot ask it for the length of a certain word there's def more to the problem than just GPT's supposed inability to determine the length of a word
-
Testing prompt modifications on complex queries
By
–
good find, def want to test this on other complex prompts to see if adding that line causes similar behavior across the board
-
LLMs as Advanced Editors: Exploring Beyond Basic Usage
By
–
yeah nothing wrong with using it the first way, LLMs serve as the primary editor for my newsletter lol but this test is mainly to highlight that most don't think to explore the next level
-

Drawing Limitations: Crayon Use Beyond Basic Line Work
By
–
this is in response to this @repligate thread @jackclarkSF argument has merit because most will only draw the line, let the crayon finish the picture, and believe that’s all it’s good for
-

The Harold and Purple Crayon Test for Prompt Engineers
By
–
the harold and the purple crayon test for determining if you are a good prompt engineer: do you draw a line and let the crayon complete the rest of the picture or do you draw the picture and let the crayon create a world?