Every time you are chatting with claude(Opus 4.5 esp. but seems to apply to all models), you quickly realize how clever claude post-training is. The attention to details, data quality. It's rare that claude will break in the middle of conversations. I think this is beyond system prompt and training dynamics/physics, and more about aura-like things: behaviors, personalities, when to push back/when to defer, and countless subtle edge cases. It's no surprise alignment people there follow closely the training. There was an invited talk at CMU this past spring on alignment and I asked why claude vibes and post-training feel different, the answer was high-level as you'd expect, but same: training, post-training, alignment, and evals teams are closed-loop. Sam Bowman (@sleepinyourhat) From everything we know so far, Opus 4.5 seems to be the best-aligned model out there in a bunch of ways. I follow the training process closely as part of my work on alignment evaluations. Here's my guess about the two things that are most responsible for making 4.5 special. 🧵 — https://nitter.net/sleepinyourhat/status/1997006353647522098#m
ETHICS
-
Pope Challenges AI Inevitability Narrative and Generated Content Issues
By
–
Interesting to see the pope point out the problems with AI-generated content while also noting the future of AI shouldn't be seen as a fait accompli. "It is a confidence that today is increasingly eroded by the paralyzing idea that its development follows an inevitable path."
-
Pope Questions AI’s Role in Common Good Distribution
By
–
Key points/questions in @Pontifex
's address today about AI's role in the world: "How can we ensure that the development of artificial intelligence truly serves the common good, and is not just used to accumulate wealth and power in the hands of a few?" -
The Blurring Line Between User and AI Prompting in 2026
By
–
The boundary between you prompting the model and the model prompting you is going to get blurry in 2026
-

Building Software That Wins Capitalism And Benefits Society
By
–
This is going to be one of the most fascinating things to figure out in the next few years, how to build software that both wins capitalism and is good for us.
-

AI Models Struggle With False Beliefs and Misconceptions Recognition
By
–
“AI needs to recognize and acknowledge false beliefs and misconceptions. That’s still a big gap in current models, even the most recent ones,” says @StanfordHAI faculty affiliate @james_y_zou on AI's current blind spots: https://
news.stanford.edu/stories/2025/1
1/ai-language-models-facts-belief-human-understanding-research
… -
Meta-Prompting: The Nuclear Option
By
–
Technique 10: Meta-Prompting (The Nuclear Option) This is what OpenAI's red team uses to break their own models and find edge cases. You ask the AI to generate the perfect prompt for itself. Template: I need to accomplish: [high-level goal] Your task:
1. Analyze what would -
Multi-Perspective Prompting Technique
By
–
Technique 9: Multi-Perspective Prompting Anthropic's Constitutional AI uses multiple viewpoints to reduce bias and improve reasoning. Template: Analyze [topic/problem] from these perspectives: [PERSPECTIVE 1: Technical Feasibility]
[specific lens] [PERSPECTIVE 2: Business -
Context Injection with Boundaries
By
–
Technique 6: Context Injection with Boundaries Anthropic engineers inject massive context but set clear boundaries on what matters. Template: [CONTEXT]
[paste your documentation, code, research paper] [FOCUS]
Only use information from CONTEXT to answer. If the answer isn't in -

Technique 5: Confidence-Weighted Prompting
By
–
Technique 5: Confidence-Weighted Prompting Google DeepMind uses this technique for high-stakes decisions. Ask the model to rate its confidence and provide alternative answers. Template: Answer this question: [question] For your answer, provide:
1. Your primary answer
2. Confidence level