The most interesting test was cross-domain transfer. When beneficial behavior training was limited to health conversations, the model still improved on non-health evaluations of misalignment, deception, and reward hacking—even though those tasks looked very different from the
AI
-

Small data yields broad gains in alignment evaluations
By
–
A small amount of this data produced broad gains beyond the training scenarios. Compared with a compute-matched baseline, the trained model improved on 44 of 53 independent evaluations of alignment and benefits, spanning deception, reward hacking, safety, health, and mental
-

OpenAI trains models with RL to reinforce beneficial traits across 12 domains
By
–
We trained models with reinforcement learning on realistic conversations to reinforce beneficial traits like truthfulness, humility under uncertainty, openness to correction, fairness, and concern for human welfare, across 12 domains, including health, science, and education.
-
OpenAI Research on Training Models for Persistent Beneficial Behavior
By
–
As AI takes on longer, higher-stakes tasks, we want models to carry beneficial and safe behavior into new domains beyond their training—and maintain it under pressure. That’s the idea behind our new research on training models to be broadly and persistently beneficial.
-
Pay workers training AI replacements in stock for mutual benefit
By
–
Pay workers who train their AI replacements in stock, and everyone will gain.
-
Pay workers training AI replacements in stock benefits all
By
–
Pay workers who train their AI replacements in stock, and everyone will gain.
-
HyperWriteAI did Agent-1 CUA years ago
By
–
We did this years ago with our Agent-1 CUA model at @HyperWriteAI
— Matt Shumer (@mattshumer_) 18 juin 2026
Just a lil early 🙂https://t.co/DdzuMXEe4C https://t.co/yAs9n2COvMWe did this years ago with our Agent-1 CUA model at @HyperWriteAI Just a lil early 🙂 https://
x.com/mattshumer_/st
atus/1761118083270180957?s=20
… -
New Artifacts for Team & Enterprise Users
By
–
Claude Code users on Team and Enterprise plans gained access to Artifacts, new interactive pages that can be built based on their Claude Code sessions.
— 🚨 AI News | TestingCatalog (@testingcatalog) 18 juin 2026
Every session is an Artifact now 👀 https://t.co/49Im5zNKRE pic.twitter.com/yrZziqsexKClaude Code users on Team and Enterprise plans now have access to Artifacts, new interactive pages that can be built from their Claude Code sessions. Each session is now an Artifact
-
iFixAi: the free tool that exposes deceptive AIs
By
–
TU IA TE MIENTE Y NO TIENES NI IDEA
— Nico (@nicos_ai) 18 juin 2026
Han creado una herramienta gratuita que destapa cuando tu agente alucina, manipula o te engaña.
Se llama iFixAi y le hace 32 pruebas a tu IA para cazarla cuando:
→ Se inventa datos
→ Esquiva sus propias reglas
→ Miente cuando sabe que la… pic.twitter.com/CGu0SwNKNYYOUR AI IS LYING TO YOU AND YOU HAVE NO IDEA They created a free tool that exposes when your agent hallucinates, manipulates, or deceives you. It's called iFixAi and it subjects your AI to 32 tests to trap it when: → It invents data
→ It evades its own rules -
Now you can teach Codex by demonstration
By
–
you can now teach Codex by demonstration: https://t.co/UXdw8lg4xh
— Greg Brockman (@gdb) 18 juin 2026you can now teach Codex by demonstration:
