We are also extremely proud of how rigorous and thorough our pre-release testing process is, including a full stack of frontier risk evaluations according to our Preparedness Framework and external red teaming. Read more safety work and evaluations in:
@lilianweng
-

OpenAI Releases o1: Advanced Reasoning Model with Improved Safety
By
–
Finally o1 is out – our first model with general reasoning capabilities. Not only it achieves impressive results on hard, scientific tasks, but also it gets significantly improved on safety and robustness. https://
openai.com/index/learning
-to-reason-with-llms/
… We found reasoning in context about safety -
Iterative AI Deployment Built on Rigorous Safety Science
By
–
Iterative deployment for maximizing AI safety learning needs to be built on top of rigorous science and process. We are learning and improving through each launch.
-
AI Security Forum Chat at Defcon 2026
By
–
Join us if you are interested for a chat at Defcon & the AI Security Forum!
-
Rule-Based Rewards for AI Safety and Capability Alignment
By
–
Rule-based rewards (RBRs) use model to provide RL signals based on a set of safety rubrics, making it easier to adapt to changing safety policies wo/ heavy dependency on human data. It also enables us to look at safety and capability in a more unified lens as a more capable
-
AI Hallucinations: How LLMs Generate Creative Misinformation
By
–
Wrote about extrinsic hallucinations during the July 4th break. https://
lilianweng.github.io/posts/2024-07-
07-hallucination/
… Here is what ChatGPT suggested as a fun tweet for the blog: Dive into the wild world of AI hallucinations! Discover how LLMs can conjure up some seriously creative (and sometimes -
ChatGPT Real-time Translation Utility During Japan Trip
By
–
I’ve started using the similar function during my Japan trip like translating my conversation with a sushi chef or teaching different types of rocks in a souvenir store. The utility is on an another level. Proud to be part of it. ❤️
— Lilian Weng (@lilianweng) 14 mai 2024
Tip: You need to interrupt the ChatGPT voice… https://t.co/pTcPubXskBI’ve started using the similar function during my Japan trip like translating my conversation with a sushi chef or teaching different types of rocks in a souvenir store. The utility is on an another level. Proud to be part of it. Tip: You need to interrupt the ChatGPT voice
-
OpenAI Hiring for Safety Systems Positions
By
–
More work coming up
& we are hiring: https://
openai.com/careers/search
?c=safety-systems
… -
Refactoring Diffusion Models: Updates on Generative AI Research
By
–
Spent some time refactoring the 2021 post on diffusion model with new content: https://
lilianweng.github.io/posts/2021-07-
11-diffusion-models/
… Then another short piece on diffusion video models: https://
lilianweng.github.io/posts/2024-04-
12-diffusion-video/
… (Yes, I had an intensive weekend) -
Data Quality and Human Factor in AI Systems
By
–
I've been thinking about data quality & human factor in the process a lot lately, so write a short post on the topic: https://
lilianweng.github.io/posts/2024-02-
05-human-data-quality/
… More: If you are into the topic, my team is hiring Research Engineer for a new sub-team Human-AI Interaction: https://
openai.com/careers/resear
ch-engineer-human-ai-interaction
…