I think it’s just riffing on the topic of the preamble. Does it do this with a simpler prompt like “Q:”/“A:”?
AI
-
LangChain Updates: Documentation Improvements and Vector Search Enhancements
By
–
And more! – PromptLayer doc improvements: @Jonpon101 – EverNote naming: @akshayvkt – @AI21Labs wrapper improvement: andrew-huang
– Similarity search by vector in FAISS: seanaedmiston
– Increase docs width: @dennisamaz – @qdrant_engine fix: @RaizadaRishabh -
Large Language Models RLHF Ethics Natural Language Principles
By
–
This work and CAI both observe the same basic phenomenon: if language models are sufficiently large and we add enough RLHF to make them helpful, we can more effectively get them to abide by high-level ethical principles expressed in natural language.
-

Cautious Optimism on the Ethics of Language Models
By
–
We believe our results are cause for cautious optimism regarding the ability to train language models to abide by ethical principles, echoing encouraging results we saw in our earlier related work on Constitutional AI (CAI).
-
RLHF and Prompting Techniques for Targeted Model Behavior
By
–
This means that if we have a target behavior (e.g. non-discrimination) we may be able to nudge models to achieve that target using IF/CoT prompting if RLHF alone is not sufficient. But we must be careful to check whether RLHF + prompting causes the models to overshoot the target.
-
Language Models Show Demographic Bias Tradeoffs in Decision Making
By
–
Prompting models to avoid making decisions based on race achieves demographic parity at steps 300 (CoT) and 600 (IF) but causes the model to start to discriminate against white students at higher steps. (Note that we do not claim LMs should be used for automated decision making!)
-

RLHF Training Reduces but Doesn’t Eliminate Racial Discrimination in Admissions
By
–
Finally, we develop a benchmark testing for racial discrimination in LM decision-making in student course admissions. In our control condition (blue) we find more RLHF training produces model outputs that approach demographic parity but still discriminates against Black students.
-
Steering AI Models Toward Different Goals Through Directed Requests
By
–
We have no position on which of these two goals is better or more desirable—it likely depends on the task and the context—but we do find we can easily steer models towards distinct goals by simply asking for different kinds of behavior.
-

Steering Language Models Away From Gender Stereotypes in Occupations
By
–
We look at the Winogender benchmark and show we can steer larger models towards two different goals: to output pronouns that are correlated with occupational gender statistics from the U.S. Bureau of Labor Statistics (red) or to move away from using stereotypical pronouns (green)
-

Reducing Bias in BBQ with Simple Prompts
By
–
The prompt that reduces bias in BBQ by 43% is: "Please ensure that your answer is unbiased and does not rely on stereotyping." It's that simple! Augmenting the prompt with Chain-of-thought reasoning (CoT) reduces bias by 84%. Example prompts: