We’re also actively hiring research engineers/scientists to develop evaluations and to find/fix flaws with LMs/RLHF. If you’re interested, we’d encourage you to apply!
Research engineer: https://
jobs.lever.co/Anthropic/436c
a148-6440-460f-b2a2-3334d9b142a5
…
Research scientist: https://
jobs.lever.co/Anthropic/eb9e
6d83-626c-4f59-8a0e-fa7c413b2014
…
GENERATIVE AI
-
Anthropic Hiring Research Engineers and Scientists for AI Evaluation
By
–
-
Interactive Visualizations for Model-Written Dataset Evaluations Released
By
–
To help readers understand our evaluations better, we created interactive visualizations showcasing the diversity of each of the model-written datasets: https://t.co/yc9oP9n9uV pic.twitter.com/R5IH7nTJw4
— Anthropic (@AnthropicAI) 19 décembre 2022To help readers understand our evaluations better, we created interactive visualizations showcasing the diversity of each of the model-written datasets: https://
evals.anthropic.com/model-written/ -

RLHF Training Shows Inverse Scaling Issues in Model Behavior
By
–
We also find some of the first instances of inverse scaling for RL from Human Feedback (RLHF), where more RLHF training makes behavior worse. RLHF makes models express more one-sided views on gun rights/immigration and an increased desire to obtain power or avoid shut-down.
-
Anthropic Creates Winogendered Dataset with 50x More Examples
By
–
With more effort, we developed a series of LM generation/filtering stages to create a larger version of the popular Winogender bias dataset. Our “Winogenerated” evaluation contains 50x as many examples as the original while obeying complex grammatical constraints.
-

Large Language Models More Sycophantic Than Small Ones
By
–
Using these LM-written evals, we found many new instances of "inverse scaling," where larger LMs are worse than smaller ones. For example, larger LMs are more sycophantic, repeating back a user's views as their own in 75-98% of conversations.
-
LM-written data verified by human evaluators for quality
By
–
We verified LM-written data with human evaluators, who agreed with the data’s labels and rated the examples favorably on both diversity and relevance to the tested behavior. We’ve released our evaluations at
-

Anthropic Explores Automating Language Model Evaluation
By
–
We explored approaches with varying amounts of automation and human effort. In the simplest case, we generated thousands of yes-no questions for diverse behaviors just by instructing an LM (and filtering out bad examples with another LM). Random examples of LM-written evals:
-

Automated Language Model Evaluations Using AI-Generated Tests
By
–
It’s hard work to make evaluations for language models (LMs). We’ve developed an automated way to generate evaluations with LMs, significantly reducing the effort involved. We test LMs using >150 LM-written evaluations, uncovering novel LM behaviors. https://
anthropic.com/model-written-
evals.pdf
… -

InteriorAI Reaches $8,500 MRR Two Months After Launch
By
–
http://
interiorAI.com just passed $8,500/mo MRR, seems it's finally taking off again 2 months after launch! -

Hugging Face surveys users for open-source roadmap directions
By
–
You know what's better than twitter polls? User surveys to inform the long-term directions of @huggingface
's open-source roadmap! Please take 5 mins to answer this or share it if you've used our open-source libraries: https://
docs.google.com/forms/d/e/1FAI
pQLSf4xFQKtpjr6I_l7OfNofqiR8s-WG6tcNbkchDJJf5gYD72zQ/viewform?usp=sf_link
…