Our Preparedness Team will drive technical work, pushing the limits of our cutting edge models to run evaluations and closely monitor risks, including during training runs. Results will be synthesized in scorecards that track model risk.
@openai
-
OpenAI releases practices for governing agentic AI systems safely
By
–
Our new research white paper identifies seven practices for keeping increasingly agentic AI systems safe and accountable as they become more common and more capable. We are providing research grants for work on a range of open questions. https://
openai.com/research/pract
ices-for-governing-agentic-ai-systems
… -
$10M ML Research Grants Launched for Community
By
–
There is lots of low-hanging fruit and many promising directions for future work. We think this is an exciting opportunity for the ML research community. To kickstart more research, we're launching $10M in grants:
-
Weak-to-Strong Generalization: Beyond RLHF for Superalignment
By
–
Naive weak supervision isn't enough—current techniques, like RLHF, won't be sufficient for future superhuman models. But we also show that it's feasible to drastically improve weak-to-strong generalization—making iterative empirical progress on a core challenge of superalignment
-

Weak Supervision Enables Large Models to Match Human-Level Performance
By
–
Large pretrained models have excellent raw capabilities—but can we elicit these fully with only weak supervision? GPT-4 supervised by ~GPT-2 recovers performance close to GPT-3.5 supervised by humans—generalizing to solve even hard problems where the weak supervisor failed!
-

Weak-to-Strong Generalization: Supervising Smarter AI Systems
By
–
In the future, humans will need to supervise AI systems much smarter than them. We study an analogy: small models supervising large models. Read the Superalignment team's first paper showing progress on a new approach, weak-to-strong generalization: https://
openai.com/research/weak-
to-strong-generalization
… -
OpenAI Announces $10M Superalignment Fast Grants Program
By
–
We're announcing, together with @ericschmidt
: Superalignment Fast Grants. $10M in grants for technical research on aligning superhuman AI systems, including weak-to-strong generalization, interpretability, scalable oversight, and more. Apply by Feb 18! -
Solving AI Alignment: A Critical Technical Challenge Ahead
By
–
Figuring out how to ensure future superhuman AI systems are aligned and safe is one of the most important unsolved technical problems in the world. But we think it is a solvable problem. There is lots of low-hanging fruit, and new researchers can make enormous contributions!
-
ChatGPT Partners with Axel Springer for Real-Time News Integration
By
–
We have formed a new global partnership with @AxelSpringer and its news products. Real-time information from @politico
, @BusinessInsider
, European properties @BILD and @welt
, and other publications will soon be available to ChatGPT users. ChatGPT’s answers to user queries will -
AI Safety Systems Team Ensures Model Reliability and Societal Benefits
By
–
Our Safety Systems team is on the frontlines of ensuring the safety and reliability of our AI models in the real world today. Learn about the team’s vision, challenges, and structure and see how you can be part of making AI safer and more beneficial for society.
