Last year, we introduced a new method for training classifiers (which stop AIs from being jailbroken to produce information about dangerous weapons). These classifiers were trained using a constitution specifying requests to which Claude should and shouldn't respond.
ETHICS
-

Claude’s Classifiers Reduce Jailbreak Success Rate to 4.4%
By
–
The classifiers reduced the jailbreak success rate from 86% to 4.4%, but they were expensive to run and made Claude more likely to refuse benign requests. We also found the system was still vulnerable to two types of attacks, shown in the figure below:
-
Anthropic’s Constitutional Classifiers Advance Jailbreak Protection
By
–
New Anthropic Research: next generation Constitutional Classifiers to protect against jailbreaks. We used novel methods, including practical application of our interpretability work, to make jailbreak protection more effective—and less costly—than ever.
-

AI Productivity Gains Require Disciplined Debt Management
By
–
Generative AI can boost developer productivity by up to 55%, but rushing deployment can pile up serious technical debt. Clear guardrails, disciplined debt management, and responsible AI training are essential to avoid costly system failures. https://
mitsmr.com/4mKf9mL @mitsmr -
Hugging Face Legal Warning Over Job Offer Data Disclosure
By
–
please Matthias do not disclose this data about our job offers at @huggingface or you'll hear from our lawyers
-
AI Productivity Gains and Unintended Consequences Analysis
By
–
Links:
Higher productivity and ROI https://
jamanetwork.com/journals/jaman
etworkopen/fullarticle/2843524
…
Unintended consequences -

Ambient Conversation AI Scribes Boost Productivity with Trade-offs
By
–
4 papers today on ambient conversation AI scribes.
They increase productivity, but……. -
Isolated tech development creates unique values and independent approaches
By
–
I’ve seen all parts of it. Good and bad. But what ended up happening there is that they had enough inward facing time to come up with everything from their own lingo, rules, themes, ideals without being influenced by the globe. Their aspirations and approaches are a bit different
-
Why AI Isn’t Delivering Expected Value
By
–
Why AI Isn't Delivering The Value You Expected
#AI #AIio #AIInnovation #ML #DataScience #Futureofwork @demishassabis @Ronald_vanLoon @TamaraMcCleary @geoffreyhinton @goodfellow_ian @jeffdean @erikbryn -
Integrating AI into Clinical Workflows: Doctors and Nurses’ Roles
By
–
Understanding the multifaceted nature of doctors' and nurses' roles is essential when integrating AI solutions into clinical workflows. The tools organisations build must be clinically sound, practical, and empathetic.
#healthcare #AI