In the blog linked below, we show real examples we found while training a recent frontier reasoning model, e.g. a model in the same class as OpenAI o1 or OpenAI o3‑mini. We found the model thinking things like, “Let’s hack,” “They don’t inspect the details,” and “We need to
ETHICS
-

Detecting Misbehavior in Frontier Reasoning Models
By
–
Detecting misbehavior in frontier reasoning models Chain-of-thought (CoT) reasoning models “think” in natural language understandable by humans. Monitoring their “thinking” has allowed us to detect misbehavior such as subverting tests in coding tasks, deceiving users, or giving
-

Schools Calling ICE on Students: Policy and Ethics Concerns
By
–
And the part where they call the cops on their own students and lay the groundwork for ICE to kidnap their own student? Which part is that? Best or nah?
-
Agentic AI Rising: Data Scientists and Developers Face Disruption
By
–
Rise of ‘#AgenticAI’: Why data scientists, software developers should be worried
-

MASK Benchmark Tests AI Honesty Under Pressure
By
–
Part of AI alignment is staying honest under pressure. Can models hold the line when pushed to lie? @scale_ai & @cais release MASK—1,000+ real-world scenarios designed to test AI honesty under pressure we'll release SEAL rankings on a private set https://
mask-benchmark.ai -
Training Helpful Harmless Assistant with Reinforcement Learning from Human Feedback
By
–
Training a Helpful and Harmless Assistant with
Reinforcement Learning from Human Feedback slides: https://
docs.google.com/presentation/d
/1hYPWiLETSK5r_Y6sU0mWbOyAz-uevNJh3Xp6BmUnGHM/edit?usp=sharing
… paper: https://
arxiv.org/abs/2204.05862 instructgpt: -
Measuring Consciousness: Computational Challenges in AI Understanding
By
–
Yeah I agree there should be some value to tracking various aspects related to consciousness that are more computational or output-based in nature as you mention. What I'm trying to convey is that I have a suspicion we'll always fall short of fully understanding / measuring the
-
xAI restricts model evaluation without consent
By
–
xAI doesn’t allow their models to be evaluated without their consent. we are working on it
-
Consciousness and Subjective Experience in AI Intelligence
By
–
I personally suspect it is going to be harder to dispel all doubts about what consciousness truly is, specifically its nature as subjective experience. Subjective would seem to imply it can't directly be verified/measured by others. Intelligence, seen as the capacity to reach
-

Deepfake Detection Workshop: Essential Skills for Synthetic Media
By
–
The deepfake detection workshop was particularly eye-opening. We created and then detected manipulated images, understanding the technical markers that reveal tampering. In today's world of "synthetic" media, these skills feel increasingly essential