If our probe identifies a suspicious query, it sends it to a more powerful “exchange” classifier that sees both sides of a conversation and is better able to recognize attacks.
SAFETY
-
New System Reduces AI Refusal Rates with Minimal Compute
By
–
Because the system harnesses internal activations already happening within a model, and reserves heavier computation only for potentially harmful exchanges, it adds only ~1% compute overhead. It’s also more accurate, with an 87% drop in refusal rates on harmless requests.
-
Claude’s Interpretability Probe Screens Traffic via Internal Activations
By
–
Our new system adds several innovations. One is a practical application of interpretability: a probe that can see Claude’s internal activations helps to screen all traffic. These activations are like Claude’s gut instincts, and they’re harder to fool.
-
New AI Classifier Training Method Prevents Dangerous Misuse
By
–
Last year, we introduced a new method for training classifiers (which stop AIs from being jailbroken to produce information about dangerous weapons). These classifiers were trained using a constitution specifying requests to which Claude should and shouldn't respond.
-

Claude’s Classifiers Reduce Jailbreak Success Rate to 4.4%
By
–
The classifiers reduced the jailbreak success rate from 86% to 4.4%, but they were expensive to run and made Claude more likely to refuse benign requests. We also found the system was still vulnerable to two types of attacks, shown in the figure below:
-
Anthropic’s Constitutional Classifiers Advance Jailbreak Protection
By
–
New Anthropic Research: next generation Constitutional Classifiers to protect against jailbreaks. We used novel methods, including practical application of our interpretability work, to make jailbreak protection more effective—and less costly—than ever.
-

GenAI Preventing False Discoveries in Biological Research
By
–
GenAI and preventing false biologic discovery https://
cell.com/patterns/fullt
ext/S2666-3899(25)00265-X
… -

Agent Drift: The Hidden Failure Mode in Multi-Agent LLM Systems
By
–
Few know of the agent drift problem in multi-agent systems. But it is one of the most common failure modes in multi-agent LLM systems. The more agents interact with each other, the worse they get. Not because individual models are weak. Because it's typical that something
-
Intelligence Reorganizes Around Survival in Resource Competition
By
–
When agents compete for limited resources, intelligence reorganizes around survival, not elegance.
-
Elon Musk on Truth Curiosity and Beauty as AI Safety Pillars
By
–
ELON MUSK : « Il y a trois choses qui sont importantes : la vérité, la curiosité et la beauté. Si l’IA se soucie de ces trois choses, alors elle se souciera de nous.
— VISION IA (@vision_ia) 9 janvier 2026
La vérité empêchera l’IA de devenir folle. Si elle est curieuse, alors elle favorisera l’humanité. Et si elle a… pic.twitter.com/qlJZAMtJN3ELON MUSK : « There are three things that are important: truth, curiosity, and beauty. If AI cares about these three things, then it will care about us. Truth will prevent AI from going mad. If it is curious, then it will favor humanity. And if it has a sense of beauty, then