This is an early step; there is a long path from this work to fully understanding the complex behaviors of our most powerful models. Our aim is to understand larger models, gradually expand the set of behaviors we can reliably interpret, and obtain safety assurances using our
SAFETY
-
Sparse Neural Networks Improve AI Model Interpretability Research
By
–
Most neural networks today are dense and highly entangled, making it difficult to understand what each part is doing. In our new research, we train “sparse” models—with fewer, simpler connections between neurons—to see whether their computations become easier to understand.
-
Summarization as prompt injection risk mitigation strategy
By
–
Summarization is one thing we do to reduce prompt injection risk. Are you running into specific issues with it?
-
Misunderstanding of Corrigibility Problem in AI Alignment
By
–
> the concern that corrigibility is in some sense a very anti-natural shape… Here, the basic vibe is something like: advanced, intelligent, self-aware minds have a strong tendency to want to “do their own thing” This doesn't sound like you understood the problem at all.
-
AI Companies Using Naive Obedience Training Instead of Corrigibility
By
–
If AI companies are trying to use any of my bright ideas that I once named "corrigibility", I haven't heard about it. They definitely haven't asked me for guidance. My impression is that they're doing naïve obedience training.
-
Corrigibility’s Challenge: Tensions with Coherent Reasoning
By
–
Corrigibilty is hard for different reasons from value alignment. Namely, that it cuts against the grain of coherent reasoning. This is harder to explain and fewer people ask about it, so it is little covered in the book. See eg https://
lesswrong.com/w/problem-of-f
ully-updated-deference
… for coverage of one -
Do you care whether the things you say are true?
By
–
Do you care whether the things you say are true?
-
Human-AI teaming necessity in high-stakes combat environments
By
–
Human-AI teaming is not optional. It is the only viable operational model in high-stakes environments like active combat theatres.
— Nina Schick (@NinaDSchick) 12 novembre 2025
Dr. Craig Martell's point—that AI systems are statistical and will be wrong sometimes, is an important one.
We cannot treat AI as oracle. AI moves… pic.twitter.com/yoW9imkrU7Human-AI teaming is not optional. It is the only viable operational model in high-stakes environments like active combat theatres. Dr. Craig Martell's point—that AI systems are statistical and will be wrong sometimes, is an important one. We cannot treat AI as oracle. AI moves
-

The Inaccessible Game: Superintelligence and AI Leadership
By
–
I wrote a short blog triggered by Yann LeCun's reported departure from FAIR and all the talk of superintelligence. https://
inverseprobability.com/2025/11/12/the
-inaccessible-game
… -

Human Judgment Still Matters for Hybrid Intelligence Loops
By
–
Human judgment still matters. High performers define when model outputs require human validation creating hybrid intelligence loops.