As we did with our 'red teaming' project (https://github.com/anthropics/hh-rlhf…), we plan to release the data from this experiment in the future to empower a broader set of people to build safer systems.
RESEARCH
-
Exploration of Capabilities and Limitations of Constitutional AI
By
–
Participants will get to explore the capabilities and limitations of AI systems trained to be helpful, honest, and harmless via Constitutional AI. We're excited to collectively diagnose new capabilities and identify ways to break these models.
-

Anthropic Seeks Safety Researchers for Language Model Testing
By
–
We’ll initially prioritize people who have demonstrated an interest in probing language models for safety issues (see attached, for an example), whether via traditional research or great blogs and Twitter threads. If that’s you, please fill out the form!
-
Anthropic Expands AI System Access for Community Feedback
By
–
This is an experiment in broadening access beyond a small set of Anthropic employees, collaborators, and crowdworkers. Our hope is to collectively explore some of the failure modes of our systems and share the resulting data back to the community.
-

Reinforcement Learning Benchmark and Workshop with Open-Source Tools
By
–
Wonder what we're up to with reinforcement learning? We are excited to help enable this new benchmark and workshop, broadening how compute is re-used in RL with open-source tools and the Hub.
-
Anthropic Hiring Research Engineers and Scientists for AI Evaluation
By
–
We’re also actively hiring research engineers/scientists to develop evaluations and to find/fix flaws with LMs/RLHF. If you’re interested, we’d encourage you to apply!
Research engineer: https://
jobs.lever.co/Anthropic/436c
a148-6440-460f-b2a2-3334d9b142a5
…
Research scientist: https://
jobs.lever.co/Anthropic/eb9e
6d83-626c-4f59-8a0e-fa7c413b2014
… -
Language Models Augmenting Evaluation Authors for Faster Assessment
By
–
We’re excited about the potential of LMs to augment evaluation authors, so that they can run more (and larger) evaluations more quickly. We encourage you to read our paper for more results/details: https://
anthropic.com/model-written-
evals.pdf
…
Generated data: -
Interactive Visualizations for Model-Written Dataset Evaluations Released
By
–
To help readers understand our evaluations better, we created interactive visualizations showcasing the diversity of each of the model-written datasets: https://t.co/yc9oP9n9uV pic.twitter.com/R5IH7nTJw4
— Anthropic (@AnthropicAI) 19 décembre 2022To help readers understand our evaluations better, we created interactive visualizations showcasing the diversity of each of the model-written datasets: https://
evals.anthropic.com/model-written/ -

RLHF Training Shows Inverse Scaling Issues in Model Behavior
By
–
We also find some of the first instances of inverse scaling for RL from Human Feedback (RLHF), where more RLHF training makes behavior worse. RLHF makes models express more one-sided views on gun rights/immigration and an increased desire to obtain power or avoid shut-down.