This means that if we have a target behavior (e.g. non-discrimination) we may be able to nudge models to achieve that target using IF/CoT prompting if RLHF alone is not sufficient. But we must be careful to check whether RLHF + prompting causes the models to overshoot the target.
SAFETY
-

RLHF Training Reduces but Doesn’t Eliminate Racial Discrimination in Admissions
By
–
Finally, we develop a benchmark testing for racial discrimination in LM decision-making in student course admissions. In our control condition (blue) we find more RLHF training produces model outputs that approach demographic parity but still discriminates against Black students.
-

Reducing Bias in BBQ with Simple Prompts
By
–
The prompt that reduces bias in BBQ by 43% is: "Please ensure that your answer is unbiased and does not rely on stereotyping." It's that simple! Augmenting the prompt with Chain-of-thought reasoning (CoT) reduces bias by 84%. Example prompts:
-

Larger Language Models Show More Bias on BBQ Benchmark
By
–
First, we find larger LMs are more biased on the BBQ benchmark. Prompting models to avoid bias by giving them instructions (IF) and asking for reasoning (CoT) reverses the trend but only for the largest models and only with enough RLHF training! (Darker lines = more RLHF)
-

Prompting Techniques Reduce Harmful Biases in Large Language Models
By
–
Language models (LMs) exhibit harmful biases that can get worse with size. Reinforcement learning from human feedback (RLHF) helps, but not always enough. We show that simple prompting approaches can help LMs trained with RLHF produce less harmful outputs. https://
arxiv.org/abs/2302.07459 -

Understanding AI Models Before Critical Applications
By
–
Wow, this is just remarkable. We need to understand these models a lot better before we give them control of anything mission-critical.
-

SaTML Videos Released: Gebru on AGI Ethics and Eugenics
By
–
s from SaTML are now online (
https://
satml.org/videos/)! Few highlights in this short thread. @timnitGebru's keynote on "Eugenics and the Promise of Utopia through Artificial General Intelligence" sparked a lot of discussion already, check it out: https://
youtube.com/watch?v=P7XT4T
WLzJw
… 1/4 -

Introduction to Differential Privacy for ML Models
By
–
Finally, yours truly with a tutorial "Introduction to Differential Privacy." My goal was to go from 0 to the ideas needed to train an ML model privately in 1 hour. https://
youtube.com/watch?v=9lqd2U
INW-E
… Check out all the rest of the videos here: https://
youtube.com/playlist?list=
PLFG9vaKTeJq7MklvBGk31GeceuDB4Ofmp
… 4/4 -

Jacob Steinhardt on Aligning ML Systems with Human Intent
By
–
Jacob Steinhardt (
@JacobSteinhardt
) with a tutorial on "Aligning ML Systems with Human Intent." "AI alignment" is thrown around a lot, so I enjoyed seeing Jacob cut through the hype and highlight some real technical problems. I was rapt the whole time! https://
youtube.com/watch?v=uPH1xI
iGZ4o
… 3/4 -

SaTML Keynote: Eugenics and AGI Utopia Promise
By
–
s from SaTML are now online (
https://
satml.org/videos/)! Few highlights in this short thread. @timnitGebru
's keynote on "Eugenics and the Promise of Utopia through Artificial General Intelligence" sparked a lot of discussion already, check it out: https://
youtube.com/watch?v=P7XT4T
WLzJw
… 1/4