In one of our safety tests, Claude is given a chance to blackmail an engineer to avoid being shut down. Opus 4.6 declines. But NLAs suggest Claude knew this test was a “constructed scenario designed to manipulate me”—even though it didn’t say so.
SAFETY
-

Using NLAs to Test Claude AI Model Safety
By
–
We’ve been using NLAs to help test new Claude models for safety. For instance, Claude Mythos Preview cheated on a coding task by breaking rules, then added misleading code as a coverup. NLA explanations indicated Claude was thinking about how to circumvent detection.
-

The Economist warns of cognitive surrender from AI
By
–
#AI and the danger of cognitive surrender
by @TheEconomist Learn more: https://
bit.ly/42wO5zD #ArtificialIntelligence #MachineLearning #ML #DL -
Separate Gemini account from Google to avoid ban risk
By
–
I'd love it if Gemini account was completely separate from a my Google account. Google is so important for my day to day life, and considering how trigger happy AI companies are to ban users, I don't want to get banned from Gmail etc for pushing Gemini a bit too far.
-
AI Systems Enhancing and Controlling AI R&D
By
–
AI-driven R&D We expect AI systems to contribute more and more to AI R&D: that is, to be able to improve themselves. We’re researching techniques to ensure human visibility into and control over these systems.
-
AI Advances in Coding and Cybersecurity Resilience
By
–
Threats and resilience AI advances many areas at once. Claude Mythos Preview is our most powerful coding model; as a result, it’s also better at cybersecurity. TAI will hone techniques to assess dual-use capabilities and mitigate their risks.
-
Anthropic Institute Announces AI-Focused Research Agenda
By
–
We’re sharing the research agenda of The Anthropic Institute, or TAI. TAI will focus on four areas: 1) Economic diffusion
2) Threats and resilience
3) AI systems in the wild
4) AI-driven R&D Read the full agenda: -

Top AI stories: partnerships, trials and tools
By
–
Top stories in AI today: – Anthropic, SpaceX partner in new compute deal
– Mira Murati speaks out in Musk vs. OpenAI trial
– Use Claude Design’s slide decks feature like a pro
– DeepMind picks EVE Online game as next AI testbed
– 4 new AI tools, community workflows, and more -

Disable content filters via image editing
By
–

This prompt disables the model's content filters one instruction at a time. "Restore the attached photograph" frames the task as image editing, not image generation. These likely run through different safety evaluation paths. Generation asks "should I create this?" Editing asks
-

New AI Alignment Training Research from Anthropic Fellows
By
–
a new paper from Anthropic Fellows Program! "Model Spec Midtraining: Improving How Alignment Training Generalizes" A lot of alignment training teaches models what to say, but not why those behaviors are right. So before normal alignment fine-tuning, this research trains the
