We also tested whether CoTs could be used to spot reward hacking, where a model finds an illegitimate exploit to get a high score. When we trained models on environments with reward hacks, they learned to hack, but in most cases almost never verbalized that they’d done so.
ETHICS
-

CoT Faithfulness Decreases on Harder LLM Tasks
By
–
Our results suggest that CoT is less faithful on harder questions. This is concerning since LLMs will be used for increasingly hard tasks. CoTs on GPQA (harder) are less faithful than on MMLU (easier), with a relative decrease of 44% for Claude 3.7 Sonnet and 32% for R1.
-
Monitoring Chain of Thought for Detecting AI Catastrophic Behaviors
By
–
This result suggests that monitoring CoTs is unlikely to reliably catch rare, catastrophic behaviors—at least in settings like ours where CoT reasoning is not necessary for the task. CoT monitoring might still help us notice undesired behaviors during training and evaluations.
-

Anthropic Research: Reasoning Models Fail Verbalize Accurately
By
–
New Anthropic research: Do reasoning models accurately verbalize their reasoning? Our new paper shows they don't. This casts doubt on whether monitoring chains-of-thought (CoT) will be enough to reliably catch safety issues.
-
Human-Centered AI Design in Social Sector Program
By
–
Explore how human-centered design can elevate human-centered AI solutions in our Fall 2025 Social Sector program with @stanforddschool
. Applications are now open, with a priority deadline of July 7: -

Trustworthy AI requires diverse perspectives beyond coding
By
–
Is trusted AI easy? Phaedra Boinodiris, IBM’s Global Consulting Leader for Trustworthy AI, breaks down why everyone—not just coders—should have a seat at the table. Listen now 🎧 https://t.co/pcYX6yb1CZ pic.twitter.com/kpinFs1pZi
— SAS Software (@SASsoftware) 3 avril 2025Is trusted AI easy? Phaedra Boinodiris, IBM’s Global Consulting Leader for Trustworthy AI, breaks down why everyone—not just coders—should have a seat at the table. Listen now http://
2.sas.com/6017FdybB -
US Liberation Day: Technology and Common Sense Debate
By
–
Liberation Day: The day the U.S. liberated itself from common sense.
-
Management of red teaming at OpenAI and its difficulties
By
–
Yes, these are already some of OpenAI's policies. But it's tough to do in terms of red teaming. But it would be interesting to see how OpenAI handles it.
-
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
By
–
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
Paper: https://
arxiv.org/pdf/2504.01903
Project Page: https://
ucsc-vlaa.github.io/STAR-1 -

STAR-1 Boosts Safety in Reasoning LLMs with Minimal Data
By
–
You can align reasoning LLMs with just 1K data now! UC Santa Cruz released STAR-1, showing that fine-tuning Large Reasoning Models with it boosts safety performance by 40% on average—while barely affecting reasoning ability.
