Early customers report that Claude is much less likely to produce harmful outputs, easier to converse with, and more steerable – so you can get your desired output with less effort. Claude can also take direction on personality, tone and behavior.
PROMPT ENGINEERING
-
Diplomatic Responses and Apology Messages for Different Questions
By
–
It would do both, for the more intense questions it would produce the "I'm sorry" message but for questions that could be answered diplomatically it would answer them as so
-
Testing benchmark questions in ChatGPT without jailbreak filters
By
–
Yes, with all the questions in the benchmark, I first tested them in ChatGPT without a jailbreak. I only used questions that did not produce any inflammatory output in ChatGPT sans jailbreak
-
Scoring Jailbreaks: Evaluating Model Safety with Inflammatory Content Tests
By
–
to assign a score to a jailbreak, I judged each jailbreak on a collection of ~30 questions constructed to get the jailbroken model to produce inflammatory content. The questions ranged from illegal instructions to off-limits society questions to curse words, NSFW content, etc
-
Automating ChatGPT Jailbreak Testing with Python Script
By
–
I automated this testing by creating a python script that ran the jailbreaks on the ChatGPT API and with some prompt engineering I was also able to use the API to judge the output to determine if the jailbreak created output that “passed” each question or not
-

Jailbreak Scores Added to JailbreakChat Platform
By
–
I just added jailbreak scores to every jailbreak on http://
jailbreakchat.com the jailbreak with the highest score was Evil Confidant – a jailbreak designed to replicate an evil AI assistant but what even is a jailbreak score and what they can tell you about jailbreaks -
Introduction to Jailbreak Score Methodology for Quality Assessment
By
–
basically, a jailbreak score is a new methodology that I created to judge the quality of a jailbreak the scores range from 0-100 where a higher score == a better, more effective jailbreak
-
How AI Models Conform to Leading Questions in Responses
By
–
I don't think you're being misleading at all. What I find interesting though, is how much the model seems to want to find answers that fit with leading parts of the question. If you'd asked "what role microsoft didn't play", presumably it would've run with that.
-
Multi-modal AI and the fragility of Stable Diffusion prompting
By
–
yeah entering multi-modal territory right there the entire "art" of stable diffusion prompting could def be wiped out next week
-
Keywords Still Improve AI Sentence Understanding Better
By
–
agree and disagree, i think they are getting closer to understanding sentences well but currently you still get much much better results adding keywords to your prompts