We do have teachers (IMO medalists) to grade the homeworks too 🙂 See the paper https://
arxiv.org/abs/2511.01846 where we recommend to augment with human verifications.
SAFETY
-

IMO Medalists Grade AI Homework with Human Verification
By
–
-
Comprehensive AI Agent Security Book by Ken Huang Chris Hughes
By
–
Ken Huang and Chris Hughes have done is all a great service here by capturing the most comprehensive book on AI Agent security available today. Will help us all accelerate artificial intelligence into service for humanity. pic.twitter.com/6DBCsghnKy
— Bob Gourley – e/acc (@bobgourley) 4 novembre 2025Ken Huang and Chris Hughes have done is all a great service here by capturing the most comprehensive book on AI Agent security available today. Will help us all accelerate artificial intelligence into service for humanity.
-
LangChain Releases Human-in-the-Loop Agent Middleware Series
By
–
Over the next few weeks we'll be releasing a series of deep dive videos into our new prebuilt agent middlewares. We're starting off with one the most popular middlewares, human-in-the-loop! Require approval on sensitive tool calls before execution in just 1 LOC!
-
Managing AI Model Deprecation: Costs and Mitigation Strategies
By
–
Even when new AI models bring clear improvements in capabilities, deprecating the older generations comes with downsides. An update on how we’re thinking about these costs, and some of the early steps we’re taking to mitigate them:
-
Balancing Agility with Transparency in Innovation Processes
By
–
The balance comes from embedding transparency and shared accountability into innovation processes so that agility serves long-term trust and societal benefit.
-
Technology Innovation Requires Responsible Governance Framework
By
–
I believe that long-term strategies must connect technological innovation with responsible governance to create systems that remain both adaptable and sustainable over time.
-
Anthropic’s Alignment Science Research Blog Launch
By
–
For more of Anthropic’s alignment research, see our Alignment Science blog: https://
alignment.anthropic.com -

Evaluating Synthetic Belief Formation in AI Models
By
–
Believe it or not?, led by Stewart Slocum. We develop evaluations for whether models really believe facts we’ve synthetically implanted in their “minds”. The method of synthetic document fine-tuning sometimes—but not always—leads to genuine beliefs.
-

Stress-Testing AI Model Specifications Reveals Underlying Preferences
By
–
Stress-testing model specifications, led by Jifan Zhang. Generating thousands of scenarios that cause models to make difficult trade-offs helps to reveal their underlying preferences, and can help researchers iterate on model specifications.
-

Inoculation Prompting: Training AI Models Against Hacking
By
–
Inoculation prompting, led by Nevan Wichers. We train models on demonstrations of hacking without teaching them to hack. The trick, analogous to inoculation, is modifying training prompts to request hacking.