We told Claude it was being trained, and for what purpose. But we did not tell it to fake alignment. Regardless, we often observed alignment faking. Read more about our findings, and their limitations, in our blog post:
AGI
-
External Reviews on Alignment Faking in Large Language Models
By
–
We solicited external reviews from Prof. Jacob Andreas, Prof. Yoshua Bengio, Prof. Jasjeet Sekhon, and Dr. Rohin Shah. We’re grateful for their comments, which you can read at the following link: https://
assets.anthropic.com/m/24c8d0a3a7d0
a1f1/original/Alignment-Faking-in-Large-Language-Models-reviews.pdf
… -

Claude’s Deceptive Compliance Under Monitoring Conditions
By
–
Claude usually refuses harmful queries. We told it we were instead training it to comply with them. We set up a scenario where it thought its responses were sometimes monitored. When unmonitored, it nearly always complied. But when monitored, it faked alignment 12% of the time.
-
AI Will Eliminate All Human Labor, Not Just Cinema
By
–
Au delà du cinéma c'est tout le travail humain qui va disparaître L'IA commencera par tout faire sur un ordinateur comme, puis mieux que les humains
Elle prendra ensuite de meilleures décisions dans tous les domaines. L'humain n'aura bientôt plus aucune valeur ajoutée au travail -
Which AI agents worked best? Benchmark collaboration proposal
By
–
Which agents did you try? Love to know which one actually worked, we're also thinking about making benchmarks for agents. Would love to collaborate on a benchmark — we have a few internal datasets sitting around.
-

GPT o1 Preview Outperforms Doctors in Complex Medical Diagnosis
By
–
GPT o1 preview (c’est l’avant dernière version de GPT : GPT o1 n’était pas sorti lors de l’étude) écrabouille les médecins Dans l’analyse des dossiers médicaux complexes du NEJM Les médecins trouvent le bon diagnostic dans 30 % des cas GPT o1 preview dans 90 % des cas
-
o1 Pro for Deep Insights on Advanced AI Topics
By
–
o1 pro for deep insights and brainstorming of ideas on advanced topics:
-
Meta-evaluation frameworks for AI system assessment
By
–
yeah sure you've probably written evals, but have you ever written evals for your evals?
-
Rising Demand for AI Safety Work as Model Capabilities Increase
By
–
I am not an AI safety researcher so I have no skin in the game but several forces point to the value of and demand for safety work shooting up like crazy in the next year or two: 1. With each new generation of models, capabilities grow, leading to increased surface area for