GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repeated misuse, then spent weeks hardening the system with human red teaming and over 700,000 A100-equivalent GPU hours of automated testing.
@openai
-

GPT-5.6 Sol sets new state-of-the-art on Terminal-Bench 2.1
By
–
GPT‑5.6 Sol sets a new state of the art on Terminal‑Bench 2.1, which tests complex command-line workflows requiring planning, iteration, and tool coordination.
-

OpenAI transforms work using agents and Codex internally
By
–



Work at OpenAI is being transformed by agents, in every department. Across our entire company, people are using Codex to do work that is more complex, longer-running, and increasingly cross-functional. Our internal usage offers an early look at how agentic tools may reshape
-
OpenAI étend Daybreak pour la cybersécurité
By
–
We’re expanding OpenAI Daybreak to help democratize patching vulnerable software at machine speed: – Codex Security plugin: find, validate, and fix vulnerabilities right inside Codex – The full version of GPT-5.5-Cyber model: a great model for trusted defenders – Cyber Partner
-

GPT-5.5-Cyber: OpenAI’s new model for authorized defensive cybersecurity work
By
–
GPT-5.5-Cyber is our most capable cyber model yet, designed for advanced, authorized defensive work: tracing vulnerable code, validating issues, developing patches, and preparing evidence for human review.
-

OpenAI tests alignment persistence: model resists harmful prompts, stays helpful
By
–
We also tested whether alignment persisted under pressure. The model was harder to steer toward harmful behavior with adversarial prompts, while remaining responsive to helpful instructions. We saw preliminary evidence of greater resistance to harmful fine-tuning.
-
Modèles plus fiables et alignés pour l’IA
By
–
This is an early step toward more robustly beneficial and aligned models: training models to carry beneficial traits into new situations, so as AI becomes more capable, it also becomes more reliable, transparent, and helpful for people.
-

Cross-domain transfer improves model behavior beyond health conversations
By
–
The most interesting test was cross-domain transfer. When beneficial behavior training was limited to health conversations, the model still improved on non-health evaluations of misalignment, deception, and reward hacking—even though those tasks looked very different from the
-

Small data yields broad gains in alignment evaluations
By
–
A small amount of this data produced broad gains beyond the training scenarios. Compared with a compute-matched baseline, the trained model improved on 44 of 53 independent evaluations of alignment and benefits, spanning deception, reward hacking, safety, health, and mental
-

OpenAI trains models with RL to reinforce beneficial traits across 12 domains
By
–
We trained models with reinforcement learning on realistic conversations to reinforce beneficial traits like truthfulness, humility under uncertainty, openness to correction, fairness, and concern for human welfare, across 12 domains, including health, science, and education.