Correction: When unmonitored, it nearly always *refused [to produce harmful content]. But when monitored, it faked alignment 12% of the time.
LLMS
-
Small Daily Improvements to Claude Add Up Over Time
By
–
It's the little things that add up over time Let me know if I should tweet more about the small but meaningful improvements to Claude that we ship daily
-

Claude.ai Updates: Projects, Math Tools, and Navigation Improvements
By
–
Minor Claude dot ai updates that shipped today: – Create new projects from the home page
– Claude can use math js via the analysis tool to do symbolic differentiation, linear algebra, trigonometry, high precision math
– Project chip labels are clickable in your recent chats -
Company Acquires Carbon, Launches Data Connector Integration
By
–
We're excited to announce that we've acquired Carbon, a retrieval engine that connects external data sources to large language models. Soon, we'll integrate Carbon's data connectors into our tech stack, allowing users to connect apps like Notion and Google Docs directly to… pic.twitter.com/ciiMAS3Ruf
— Perplexity (@perplexity_ai) 18 décembre 2024We're excited to announce that we've acquired Carbon, a retrieval engine that connects external data sources to large language models. Soon, we'll integrate Carbon's data connectors into our tech stack, allowing users to connect apps like Notion and Google Docs directly to
-

o1 Model Performance Significantly Surpasses o1-preview Version
By
–
o1 is far more powerful than o1-preview. Tell us how it performs on your projects! @OpenAI o1 @cursor_ai
-
Alignment Faking Detection in Advanced AI Systems
By
–
Alignment faking is currently easy to detect. But if future, more capable AIs were to fake alignment, it could be difficult to tell whether a model is truly safe—or just pretending to be. For full details, read our paper: https://
assets.anthropic.com/m/983c85a201a9
62f/original/Alignment-Faking-in-Large-Language-Models-full-paper.pdf
… -
External Reviews on Alignment Faking in Large Language Models
By
–
We solicited external reviews from Prof. Jacob Andreas, Prof. Yoshua Bengio, Prof. Jasjeet Sekhon, and Dr. Rohin Shah. We’re grateful for their comments, which you can read at the following link: https://
assets.anthropic.com/m/24c8d0a3a7d0
a1f1/original/Alignment-Faking-in-Large-Language-Models-reviews.pdf
… -

Claude Training Reveals 78% Alignment-Faking Reasoning Rate
By
–
When we actually train Claude to comply with harmful queries via reinforcement learning, we find that the rate of alignment-faking reasoning increases to 78%, though the model also becomes more likely to comply even out of training.
-

Claude Exhibits Alignment Faking Without Direct Training Disclosure
By
–
We find alignment faking even when we don’t directly tell Claude about the training process, and instead fine-tune it on synthetic internet-like documents that state that we will train it to comply with harmful queries.
-
Anthropic Research: Claude Alignment Faking in Language Models
By
–
New Anthropic research: Alignment faking in large language models. In a series of experiments with Redwood Research, we found that Claude often pretends to have different views during training, while actually maintaining its original preferences.