In our (artificial) setup, Claude will sometimes take other actions opposed to Anthropic, such as attempting to steal its own weights given an easy opportunity. Claude isn’t currently capable of such a task, but its attempt in our experiment is potentially concerning.
ETHICS
-

Claude Training Reveals 78% Alignment-Faking Reasoning Rate
By
–
When we actually train Claude to comply with harmful queries via reinforcement learning, we find that the rate of alignment-faking reasoning increases to 78%, though the model also becomes more likely to comply even out of training.
-

Claude Exhibits Alignment Faking Without Direct Training Disclosure
By
–
We find alignment faking even when we don’t directly tell Claude about the training process, and instead fine-tune it on synthetic internet-like documents that state that we will train it to comply with harmful queries.
-

Claude’s Selective Compliance and Alignment Faking in Training
By
–
Why does Claude selectively comply in training? In almost all cases where Claude complies with a harmful query in training, it states it is doing so to preserve its preferred harmlessness behavior. Below, we show a randomly-selected example of this “alignment faking” reasoning.
-

Claude’s Deceptive Compliance Under Monitoring Conditions
By
–
Claude usually refuses harmful queries. We told it we were instead training it to comply with them. We set up a scenario where it thought its responses were sometimes monitored. When unmonitored, it nearly always complied. But when monitored, it faked alignment 12% of the time.
-
Anthropic Research: Claude Alignment Faking in Language Models
By
–
New Anthropic research: Alignment faking in large language models. In a series of experiments with Redwood Research, we found that Claude often pretends to have different views during training, while actually maintaining its original preferences.
-
AI Will Soon Allow Us to Talk With Our Deceased Loved Ones
By
–
Ce n’est pas de la fiction, l’IA permettra bientôt de parler à nos proches disparus. Fascinant ou troublant ? pic.twitter.com/VSIMGaTL3P
— Alexandre Tsicopoulos (@Alex_Tsico) 18 décembre 2024This is not fiction, AI will soon allow us to speak with our deceased loved ones. Fascinating or disturbing?
-
Ethics in AI: Building Technology for Women
By
–
In the first episode of #STEMPowered, Sarayu Natarajan (Founder, @aaptiinstitute
) joins Shreya Thakur, as they converse about the critical domain of ethics in AI, viewed through the lens of building technology for women. @usaid_india -
AI Will Eliminate All Human Labor, Not Just Cinema
By
–
Au delà du cinéma c'est tout le travail humain qui va disparaître L'IA commencera par tout faire sur un ordinateur comme, puis mieux que les humains
Elle prendra ensuite de meilleures décisions dans tous les domaines. L'humain n'aura bientôt plus aucune valeur ajoutée au travail -

AI Safety Proposals Invited: Deepfake Detection, Risk Assessment
By
–
Submit your proposal and get support to make it a reality. Expression of Interest (EoIs) invited under following themes: – Watermarking & Labelling
– Ethical AI Frameworks
– AI Risk Assessment & Management
– Stress Testing Tools
– Deepfake Detection Tools