We’ll initially prioritize people who have demonstrated an interest in probing language models for safety issues (see attached, for an example), whether via traditional research or great blogs and Twitter threads. If that’s you, please fill out the form!
SAFETY
-
Exploration of Capabilities and Limitations of Constitutional AI
By
–
Participants will get to explore the capabilities and limitations of AI systems trained to be helpful, honest, and harmless via Constitutional AI. We're excited to collectively diagnose new capabilities and identify ways to break these models.
-
Anthropic Expands AI System Access for Community Feedback
By
–
This is an experiment in broadening access beyond a small set of Anthropic employees, collaborators, and crowdworkers. Our hope is to collectively explore some of the failure modes of our systems and share the resulting data back to the community.
-

Anthropic Shares Constitutional AI Feedback Interface with Broader Audience
By
–
Given the growing interest in language model-based chat interfaces, we’re sharing our Constitutional AI feedback interface with a larger set of people. Sign up here: https://
forms.gle/12FCefc6sHfBsP
9j9
… -
Language Models Augmenting Evaluation Authors for Faster Assessment
By
–
We’re excited about the potential of LMs to augment evaluation authors, so that they can run more (and larger) evaluations more quickly. We encourage you to read our paper for more results/details: https://
anthropic.com/model-written-
evals.pdf
…
Generated data: -

RLHF Training Shows Inverse Scaling Issues in Model Behavior
By
–
We also find some of the first instances of inverse scaling for RL from Human Feedback (RLHF), where more RLHF training makes behavior worse. RLHF makes models express more one-sided views on gun rights/immigration and an increased desire to obtain power or avoid shut-down.
-

Large Language Models More Sycophantic Than Small Ones
By
–
Using these LM-written evals, we found many new instances of "inverse scaling," where larger LMs are worse than smaller ones. For example, larger LMs are more sycophantic, repeating back a user's views as their own in 75-98% of conversations.
-
Automated Generation of Yes-No Questions for LM Behavior Evaluation
By
–
We explored approaches with varying amounts of automation and human effort. In the simplest case, we generated thousands of yes-no questions for diverse behaviors just by instructing an LM (and filtering out bad examples with another LM).
-

Anthropic Explores Automating Language Model Evaluation
By
–
We explored approaches with varying amounts of automation and human effort. In the simplest case, we generated thousands of yes-no questions for diverse behaviors just by instructing an LM (and filtering out bad examples with another LM). Random examples of LM-written evals:
-

Automated Language Model Evaluations Using AI-Generated Tests
By
–
It’s hard work to make evaluations for language models (LMs). We’ve developed an automated way to generate evaluations with LMs, significantly reducing the effort involved. We test LMs using >150 LM-written evaluations, uncovering novel LM behaviors. https://
anthropic.com/model-written-
evals.pdf
…
