this is going to take continual iteration–and lots and lots of societal input–to get right. to find the right balance, we will likely overcorrect several times, and find new edges in the technology. we appreciate the patience and good faith as we get to a better place!
ETHICS
-
AI System Behavior: Reducing Bias, Customization, and Public Input
By
–
our current thoughts on hard questions about how AI systems should behave: 1) less biased defaults, 2) lots of user customization within very broad bounds, 3) public input on bounds and defaults
-
ChatGPT Behavior Control: Future Governance and User Input
By
–
How ChatGPT’s behaviors are determined, and how we think it should work in the future — including thoughts on giving users much more control & early ideas around public input: https://
openai.com/blog/how-shoul
d-ai-systems-behave/
… -
ChatGPT Alignment Improvements and User Control Expansion
By
–
Information on ChatGPT’s alignment, plans to improve it, giving users more control, and early thoughts on public input:
-

ChatGPT Explanation: Impressive But Not Revolutionary
By
–
@stephen_wolfram gives a necessarily long but very good explanation of what and how chatGPT works. My view: calm down people. There are important things there but not the shiny "world has changed for ever" techno-salvation that many people pining for.
-
Large Language Models RLHF Ethics Natural Language Principles
By
–
This work and CAI both observe the same basic phenomenon: if language models are sufficiently large and we add enough RLHF to make them helpful, we can more effectively get them to abide by high-level ethical principles expressed in natural language.
-

Cautious Optimism on the Ethics of Language Models
By
–
We believe our results are cause for cautious optimism regarding the ability to train language models to abide by ethical principles, echoing encouraging results we saw in our earlier related work on Constitutional AI (CAI).
-
RLHF and Prompting Techniques for Targeted Model Behavior
By
–
This means that if we have a target behavior (e.g. non-discrimination) we may be able to nudge models to achieve that target using IF/CoT prompting if RLHF alone is not sufficient. But we must be careful to check whether RLHF + prompting causes the models to overshoot the target.
-
Language Models Show Demographic Bias Tradeoffs in Decision Making
By
–
Prompting models to avoid making decisions based on race achieves demographic parity at steps 300 (CoT) and 600 (IF) but causes the model to start to discriminate against white students at higher steps. (Note that we do not claim LMs should be used for automated decision making!)
-

RLHF Training Reduces but Doesn’t Eliminate Racial Discrimination in Admissions
By
–
Finally, we develop a benchmark testing for racial discrimination in LM decision-making in student course admissions. In our control condition (blue) we find more RLHF training produces model outputs that approach demographic parity but still discriminates against Black students.