We've made progress on the AI safety problem of detecting and reducing "scheming": – Created evaluation environments to detect scheming
– Observed current models scheming in controlled settings
– Found deliberative alignment (
https://
openai.com/index/delibera
tive-alignment/
…) decreases scheming rates
Detecting and Reducing AI Scheming Behavior Through Deliberative Alignment
By
–
