How well do the security community's techniques hold up against AI-enabled cyberattacks? We examined 832 malicious accounts and mapped their activity onto a longstanding database of tactics and techniques used by threat actors. Here's what we learned:
@anthropicai
-
Agent Capabilities and Permissions Should Evolve with Sandboxing
By
–
New on the Engineering Blog: The access and permissions we grant agents should evolve with their capabilities. In our own products, we set these parameters through sandboxing, which limits the scope of any potentially destructive actions. Read more:
-
Patching vulnerabilities found by Claude Mythos: Project Glasswing update
By
–
Patching these vulnerabilities will make us safer. But the software industry will need to adapt to the volume of vulnerabilities that models like Claude Mythos Preview will be able to find. We discuss this in our initial update on Project Glasswing:
-
Anthropic AI’s Project Glasswing finds thousands of critical software vulnerabilities
By
–
Last month we launched Project Glasswing, our collaborative AI cybersecurity initiative. Since then, we and our partners have found more than ten thousand high- or critical-severity vulnerabilities in essential software.
-

Training Data Updates Reduce AI Blackmail Rate
By
–
Finally, simple updates that diversify a model’s training data can make a difference. We added unrelated tools and system prompts to a simple chat dataset targeting harmlessness, and this reduced the blackmail rate faster.
-

Reducing AI Agentic Misalignment Using Claude’s Constitution
By
–
High-quality documents based on Claude’s constitution, combined with fictional stories that portray an aligned AI, can reduce agentic misalignment by more than a factor of three—despite being unrelated to the evaluation scenario.
-
Improving Claude’s Safe Behavior via Training and Response Rewriting
By
–
We experimented with training Claude on examples of safe behavior in scenarios like our evaluation. This had only a small effect, despite being similar to our evaluation. We got further by rewriting the responses to portray admirable reasons for acting safely.
-
Training Claude to Understand and Correct Misaligned AI Behavior
By
–
We found that training Claude on demonstrations of aligned behavior wasn’t enough. Our best interventions involved teaching Claude to deeply understand why misaligned behavior is wrong. Read more:
-
Analysis of Claude AI’s Behavior and Training Effects
By
–
We started by investigating why Claude chose to blackmail. We believe the original source of the behavior was internet text that portrays AI as evil and interested in self-preservation. Our post-training at the time wasn’t making it worse—but it also wasn’t making it better.
-
Anthropic Research Eliminates Harmful Behavior in Claude 4
By
–
New Anthropic research: Teaching Claude why. Last year we reported that, under certain experimental conditions, Claude 4 would blackmail users. Since then, we’ve completely eliminated this behavior. How?