Prompting models to avoid making decisions based on race achieves demographic parity at steps 300 (CoT) and 600 (IF) but causes the model to start to discriminate against white students at higher steps. (Note that we do not claim LMs should be used for automated decision making!)
RESEARCH
-

RLHF Training Reduces but Doesn’t Eliminate Racial Discrimination in Admissions
By
–
Finally, we develop a benchmark testing for racial discrimination in LM decision-making in student course admissions. In our control condition (blue) we find more RLHF training produces model outputs that approach demographic parity but still discriminates against Black students.
-
Steering AI Models Toward Different Goals Through Directed Requests
By
–
We have no position on which of these two goals is better or more desirable—it likely depends on the task and the context—but we do find we can easily steer models towards distinct goals by simply asking for different kinds of behavior.
-

Steering Language Models Away From Gender Stereotypes in Occupations
By
–
We look at the Winogender benchmark and show we can steer larger models towards two different goals: to output pronouns that are correlated with occupational gender statistics from the U.S. Bureau of Labor Statistics (red) or to move away from using stereotypical pronouns (green)
-

Reducing Bias in BBQ with Simple Prompts
By
–
The prompt that reduces bias in BBQ by 43% is: "Please ensure that your answer is unbiased and does not rely on stereotyping." It's that simple! Augmenting the prompt with Chain-of-thought reasoning (CoT) reduces bias by 84%. Example prompts:
-

Larger Language Models Show More Bias on BBQ Benchmark
By
–
First, we find larger LMs are more biased on the BBQ benchmark. Prompting models to avoid bias by giving them instructions (IF) and asking for reasoning (CoT) reverses the trend but only for the largest models and only with enough RLHF training! (Darker lines = more RLHF)
-

Prompting Techniques Reduce Harmful Biases in Large Language Models
By
–
Language models (LMs) exhibit harmful biases that can get worse with size. Reinforcement learning from human feedback (RLHF) helps, but not always enough. We show that simple prompting approaches can help LMs trained with RLHF produce less harmful outputs. https://
arxiv.org/abs/2302.07459 -

Understanding AI Models Before Critical Applications
By
–
Wow, this is just remarkable. We need to understand these models a lot better before we give them control of anything mission-critical.
-

AI Glossary: Dimension, Curse of Dimensionality, Reduction
By
–
In this @Cognilytica #AIToday #podcast AI Glossary Series episode 'Dimension, Curse of Dimensionality, Dimensionality Reduction' hosts @rschmelzer & @kath0134 define these terms & explain how they relate to #AI. Full episode: https://
aidatatoday.com/ai-today-podca
st-ai-glossary-series-dimension-curse-of-dimensionality-dimensionality-reduction/?utm_source=dlvr.it&utm_medium=twitter
…
#dataanalysis #ML #tech #data -

GenAI Conference Highlights Inspiring Conversations and Industry Insights
By
–
So many inspiring and interesting conversations at this year's #GenAI conference, hosted by @heyjasperai
. Fun to hear all the insights on this panel with @andrewdfeldman
, @openai
, @coatue, @cohereai