Happy to see researchers questioning whether RLHF is sufficient for “alignment”. I like this perspective: https://
researchgate.net/publication/22
8764857_That_special_something_Dennett_on_the_making_of_minds_and_selves
… of @danieldennett
. But ethical questions abound as it involves ego, TOM, morality, compassion, empathy…
SAFETY
-

RLHF Limitations and Ethical Alignment in AI Systems
By
–
-
Treating AI Agents Kindly Today Shapes Tomorrow’s AGIs
By
–
I've noticed that people are incredibly polite to ChatGPT, even thanking it for good advice. While this might seem like anthropomorphization, it's also a precaution in case future AGIs reflect how we treated their predecessors. Be kind to our AI agents – they're people too!
-
Automated Testing for Social Bias in Large Language Models
By
–
Testing Language Models at scale for social bias is challenging. We build automated testing: test sentences are automatically generated given the bias dimensions. This allows getting statistically meaningful measures for bias as opposed to a small number of hand-written templates
-
Why Chatbots Malfunction: Understanding AI Limitations and Safeguards
By
–
Let's get real. This is why chatbots sometimes act weird and spout nonsense — and why companies can't completely stop it:
-

Google penalized for chatbot error, Microsoft rewarded despite threats
By
–
Google: Advertise an unreleased chatbot, makes a minor factual error about exoplanet photos. Lose $100,000,000,000.
— Gautam Kamath (@thegautamkamath) 16 février 2023
Microsoft: Release a chatbot that threatens to hunt down and murder users, in some cases for correcting its factual errors. Stock rises.
🤷 https://t.co/qvzYA4eX6CGoogle: Advertise an unreleased chatbot, makes a minor factual error about exoplanet photos. Lose $100,000,000,000. Microsoft: Release a chatbot that threatens to hunt down and murder users, in some cases for correcting its factual errors. Stock rises.
-
Balancing AI Development Through Iteration and Societal Input
By
–
this is going to take continual iteration–and lots and lots of societal input–to get right. to find the right balance, we will likely overcorrect several times, and find new edges in the technology. we appreciate the patience and good faith as we get to a better place!
-
ChatGPT Behavior Control: Future Governance and User Input
By
–
How ChatGPT’s behaviors are determined, and how we think it should work in the future — including thoughts on giving users much more control & early ideas around public input: https://
openai.com/blog/how-shoul
d-ai-systems-behave/
… -
ChatGPT Alignment Improvements and User Control Expansion
By
–
Information on ChatGPT’s alignment, plans to improve it, giving users more control, and early thoughts on public input:
-
Large Language Models RLHF Ethics Natural Language Principles
By
–
This work and CAI both observe the same basic phenomenon: if language models are sufficiently large and we add enough RLHF to make them helpful, we can more effectively get them to abide by high-level ethical principles expressed in natural language.
-

Cautious Optimism on the Ethics of Language Models
By
–
We believe our results are cause for cautious optimism regarding the ability to train language models to abide by ethical principles, echoing encouraging results we saw in our earlier related work on Constitutional AI (CAI).