I didn’t come up with the core idea, but I wrote a modified stronger version and showed it worked on many frontier models. Earlier examples were closer to yours — the riddle, but omitting the father’s death. Mine goes further, explicitly stating the surgeon is the father.
ETHICS
-

Stanford HAI Welcomes Six Distinguished Senior Fellows
By
–
Big news! Join us in welcoming Susan Athey, Michael Bernstein, Angèle Christin, Mykel Kochenderfer, Dorsa Sadigh, and Melissa Valentine as our newest senior fellows. These esteemed scholars bring unmatched expertise and research to advance our mission. https://
hai.stanford.edu/news/stanford-
hai-welcomes-six-distinguished-scholars-senior-fellows
… -
Test Jailbreaks to Improve AI System Security
By
–
Try your favorite jailbreaks on it and help us better secure powerful AI systems: https://
claude.ai/constitutional
-classifiers
… -
Anthropic Releases AI Safety Demo App for Red-Teaming
By
–
At Anthropic, we're preparing for the arrival of powerful AI systems. Based on our latest research on Constitutional Classifiers, we've developed a demo app to test new safety techniques.
— Alex Albert (@alexalbert__) 3 février 2025
We want you to help us red-team the app – so far no one has been able to crack the… pic.twitter.com/CpwnqU1tAFAt Anthropic, we're preparing for the arrival of powerful AI systems. Based on our latest research on Constitutional Classifiers, we've developed a demo app to test new safety techniques. We want you to help us red-team the app – so far no one has been able to crack the
-
Anthropic’s Jailbreak Demo Tested
By
–
Anthropic just released a demo system where users can try to jailbreak their system My Monday evening is wasted now. Let's see if o3-mini-high can do anything there
-
Anthropic Hiring for Safeguards Research Team Roles
By
–
If you're interested in working on similar topics, you can apply for a role on Anthropic's Safeguards Research Team:
-
Constitutional Classifiers Demo: Security Challenge for AI Safety
By
–
Can you do better than our red teamers? We’ve made a demo system protected by Constitutional Classifiers. We challenge you to jailbreak it to help us make our defenses even stronger. Try the demo: http://
claude.ai/constitutional
-classifiers
… -
Constitutional Classifiers: Flexible AI Defense Against Novel Attacks
By
–
Constitutional Classifiers aren’t perfect. We recommend using other complementary defenses, such as rapid-response techniques. Nevertheless, our method is flexible, and the constitution can quickly be adapted to cover novel attacks. Read the full paper:
-
AI System Passes Rigorous Jailbreak Testing After Red Teaming
By
–
We challenged jailbreakers to try to break a prototype version of the system to test its robustness. After thousands of hours of red teaming, not one participant found a reliable jailbreak that extracted detailed information across a set of 10 harmful questions.
-

Constitutional Classifiers Reduce Jailbreak Effectiveness with Compute Trade-offs
By
–
In an experiment with synthetic jailbreaks, Constitutional Classifiers dramatically reduced jailbreak effectiveness. They increased refusal rates by a small amount (+0.4%) and increased compute overhead (+24%). We're working on reducing these costs.
