I ran an AI safety conference. And I look for answers to potential downfalls. AI tech will save many lives and make our economy way more productive. So I start there.
SAFETY
-
Chain-of-Thought Monitorability for AI Safety
By
–
New work on evaluating the quality of chain-of-thought monitorability. Chain-of-thought monitorability is a very encouraging opportunity for safety and alignment, making it easy to see what models are thinking:
-
Chain-of-thought monitoring as complementary mechanistic interpretability approach
By
–
We view chain-of-thought monitoring as complementary to mechanistic interpretability, not as a replacement for it. Because we believe that chain-of-thought monitoring is incredibly useful as a window into a model’s brain and could be a loadbearing layer in a scalable control
-

Chain-of-Thought Monitoring for Better Model Oversight
By
–
Monitoring a model’s chain-of-thought is far more effective than watching only its actions or final answers. The more a model “thinks” (longer CoTs), the easier it is to spot issues.
-
RL Reasoning Tradeoff: Monitorability vs Inference Compute
By
–
RL at today’s frontier doesn’t seem to wreck monitorability and can help early reasoning steps. But there’s a tradeoff: smaller models run with higher reasoning effort can be easier to monitor at similar capability — at the cost of extra inference compute (a “monitorability
-
Measuring Chain-of-Thought Monitorability in AI Models
By
–
To preserve chain-of-thought (CoT) monitorability, we must be able to measure it. We built a framework + evaluation suite to measure CoT monitorability — 13 evaluations across 24 environments — so that we can actually tell when models verbalize targeted aspects of their
-
OpenAI Model Spec: Intended Behavior for AI Models
By
–
The Model Spec — intended behavior for the models that power OpenAI’s products:
-
Claude Ensures Empathetic Honest Support in AI Conversations
By
–
People use AI for a wide variety of reasons, including emotional support. Below, we share the efforts we’ve taken to ensure that Claude handles these conversations both empathetically and honestly.
-
Google Adds AI Video Transparency Tools to Gemini App
By
–
Google's expanding its content transparency tools to let people analyze AI-generated videos in the Gemini app to check for the SynthID watermark.
-
Robot Safety Concerns and Human Harm Prevention Timeline
By
–
Yes but not the present. 2028 mostly. Robots can still accidentally harm humans.