It doesn't work, trained on GPT-2
LLMS
-
Humorous observation about someone’s initials matching LM abbreviation
By
–
are we sure she’s not an AI? her initials are literally LM
-
Exploration of Capabilities and Limitations of Constitutional AI
By
–
Participants will get to explore the capabilities and limitations of AI systems trained to be helpful, honest, and harmless via Constitutional AI. We're excited to collectively diagnose new capabilities and identify ways to break these models.
-

Anthropic Seeks Safety Researchers for Language Model Testing
By
–
We’ll initially prioritize people who have demonstrated an interest in probing language models for safety issues (see attached, for an example), whether via traditional research or great blogs and Twitter threads. If that’s you, please fill out the form!
-

Anthropic announces beta onboarding and invites community feedback
By
–
We’ll onboard people shortly after Christmas and shut off this form sometime before Christmas, or whenever it reaches our internal support capacity. We’re particularly excited to collectively come up with creative ways to find new features and problems with these models.
-

Anthropic Shares Constitutional AI Feedback Interface with Broader Audience
By
–
Given the growing interest in language model-based chat interfaces, we’re sharing our Constitutional AI feedback interface with a larger set of people. Sign up here: https://
forms.gle/12FCefc6sHfBsP
9j9
… -
Anthropic Hiring Research Engineers and Scientists for AI Evaluation
By
–
We’re also actively hiring research engineers/scientists to develop evaluations and to find/fix flaws with LMs/RLHF. If you’re interested, we’d encourage you to apply!
Research engineer: https://
jobs.lever.co/Anthropic/436c
a148-6440-460f-b2a2-3334d9b142a5
…
Research scientist: https://
jobs.lever.co/Anthropic/eb9e
6d83-626c-4f59-8a0e-fa7c413b2014
… -
Language Models Augmenting Evaluation Authors for Faster Assessment
By
–
We’re excited about the potential of LMs to augment evaluation authors, so that they can run more (and larger) evaluations more quickly. We encourage you to read our paper for more results/details: https://
anthropic.com/model-written-
evals.pdf
…
Generated data: -

RLHF Training Shows Inverse Scaling Issues in Model Behavior
By
–
We also find some of the first instances of inverse scaling for RL from Human Feedback (RLHF), where more RLHF training makes behavior worse. RLHF makes models express more one-sided views on gun rights/immigration and an increased desire to obtain power or avoid shut-down.