There are papers that show training AI on "evil" data results in general misalignment, so it is nice to know the opposite is true and that beneficial RL data in one field leads to more aligned models across a range of tasks.
SAFETY
-
Open Weights: Control Over Capability Even When Trailing
By
–
Open weights matter even when they trail the frontier exactly because of this, you're buying control, not just capability. weights matter even when they trail the frontier exactly because of this, you're buying control, not just capability.
-
Ineffective a posteriori API guardrails for cutting-edge models
By
–
Let's face the truth: a posteriori API guardrails are not the appropriate safety tool for cutting-edge models. They do not eliminate dangerous capabilities. They simply hide them behind a fragile interface that can be easily
-

OpenAI tests alignment persistence: model resists harmful prompts, stays helpful
By
–
We also tested whether alignment persisted under pressure. The model was harder to steer toward harmful behavior with adversarial prompts, while remaining responsive to helpful instructions. We saw preliminary evidence of greater resistance to harmful fine-tuning.
-
Modèles plus fiables et alignés pour l’IA
By
–
This is an early step toward more robustly beneficial and aligned models: training models to carry beneficial traits into new situations, so as AI becomes more capable, it also becomes more reliable, transparent, and helpful for people.
-

Cross-domain transfer improves model behavior beyond health conversations
By
–
The most interesting test was cross-domain transfer. When beneficial behavior training was limited to health conversations, the model still improved on non-health evaluations of misalignment, deception, and reward hacking—even though those tasks looked very different from the
-

Small data yields broad gains in alignment evaluations
By
–
A small amount of this data produced broad gains beyond the training scenarios. Compared with a compute-matched baseline, the trained model improved on 44 of 53 independent evaluations of alignment and benefits, spanning deception, reward hacking, safety, health, and mental
-

OpenAI trains models with RL to reinforce beneficial traits across 12 domains
By
–
We trained models with reinforcement learning on realistic conversations to reinforce beneficial traits like truthfulness, humility under uncertainty, openness to correction, fairness, and concern for human welfare, across 12 domains, including health, science, and education.
-
OpenAI Research on Training Models for Persistent Beneficial Behavior
By
–
As AI takes on longer, higher-stakes tasks, we want models to carry beneficial and safe behavior into new domains beyond their training—and maintain it under pressure. That’s the idea behind our new research on training models to be broadly and persistently beneficial.
-
iFixAi: the free tool that exposes deceptive AIs
By
–
TU IA TE MIENTE Y NO TIENES NI IDEA
— Nico (@nicos_ai) 18 juin 2026
Han creado una herramienta gratuita que destapa cuando tu agente alucina, manipula o te engaña.
Se llama iFixAi y le hace 32 pruebas a tu IA para cazarla cuando:
→ Se inventa datos
→ Esquiva sus propias reglas
→ Miente cuando sabe que la… pic.twitter.com/CGu0SwNKNYYOUR AI IS LYING TO YOU AND YOU HAVE NO IDEA They created a free tool that exposes when your agent hallucinates, manipulates, or deceives you. It's called iFixAi and it subjects your AI to 32 tests to trap it when: → It invents data
→ It evades its own rules
