As AI takes on longer, higher-stakes tasks, we want models to carry beneficial and safe behavior into new domains beyond their training—and maintain it under pressure. That’s the idea behind our new research on training models to be broadly and persistently beneficial.
@openai
-
AGI to improve human health through better ChatGPT
By
–
Improving human health will be one of the most personal, tangible impacts of AGI. As our models continue to improve, our goal is to make ChatGPT more accurate, more useful, and more impactful in those moments — and to keep bringing that progress to more people.
-
OpenAI collaborates with global physicians to improve AI models
By
–
To improve our models, we collaborate with a global network of hundreds of physicians across 60 countries, 49 languages, and 26 specialties. Their feedback helps us identify where responses miss important context, sound overly confident, need clearer next steps, or should more
-
GPT-5.5 Instant matches frontier thinking models for health questions
By
–
GPT-5.5 Instant is now on par with our frontier Thinking models for health-related questions.
— OpenAI (@OpenAI) 18 juin 2026
Every week, more than 230 million people turn to ChatGPT with health and wellness questions, and GPT-5.5 Instant is better at recognizing when urgent care may be needed, asking for… pic.twitter.com/WDDzOIV3mgGPT-5.5 Instant is now on par with our frontier Thinking models for health-related questions. Every week, more than 230 million people turn to ChatGPT with health and wellness questions, and GPT-5.5 Instant is better at recognizing when urgent care may be needed, asking for
-
LifeSciBench: A Foundation for Realistic AI Evaluation in Life Sciences
By
–
LifeSciBench is a foundation for more realistic evaluation, targeted improvements, and continued partnership with the life sciences community—helping the field measure progress, identify gaps, and improve AI together for the benefit of everyone.
-

OpenAI introduces LifeSciBench benchmark for life science research
By
–
Introducing LifeSciBench, a benchmark for measuring and improving how well AI supports real-world life science research. Developed with 173 scientists from biotechnology and pharmaceutical research, LifeSciBench includes 750 expert-authored tasks across seven biological research
-

LifeSciBench tests reasoning; GPT‑Rosalind outperforms GPT‑5.5
By
–
Benchmarks often test biological knowledge or narrow skills. The tasks in LifeSciBench test whether models can reason from evidence, work with scientific artifacts, handle uncertainty, and make useful decisions under real-world constraints. GPT‑Rosalind scores above GPT‑5.5
-
Frontier models support scientific research loop in 2.5 months
By
–
The full process took about 2.5 months, plus another half month for human chemists to write up the results. This is an early example of frontier models supporting more of the scientific research loop: reviewing studies, proposing hypotheses, designing experiments, interpreting
-
GPT-5.4 and human chemists collaborate on full research workflow
By
–
GPT-5.4 reviewed scientific literature, generated and ranked research proposals, helped design experiments, analyzed results, and proposed follow-up studies. Human chemists steered the work, selected proposals for testing, and validated the final result.
-
GPT-5.4 and Maria AI improve drug discovery reaction
By
–
GPT-5.4 helped drive a medicinal chemistry project from literature review to a validated experimental result.
— OpenAI (@OpenAI) 17 juin 2026
Paired with https://t.co/gcDaph8b2B’s Maria AI and specialized lab, the model proposed an unexpected way to improve a widely used reaction in drug discovery. pic.twitter.com/KmyBlHLX8yGPT-5.4 helped drive a medicinal chemistry project from literature review to a validated experimental result. Paired with http://
Molecule.one’s Maria AI and specialized lab, the model proposed an unexpected way to improve a widely used reaction in drug discovery.