Using these LM-written evals, we found many new instances of "inverse scaling," where larger LMs are worse than smaller ones. For example, larger LMs are more sycophantic, repeating back a user's views as their own in 75-98% of conversations.
RESEARCH
-
LM-written data verified by human evaluators for quality
By
–
We verified LM-written data with human evaluators, who agreed with the data’s labels and rated the examples favorably on both diversity and relevance to the tested behavior. We’ve released our evaluations at
-
Anthropic Creates Winogendered Dataset with 50x More Examples
By
–
With more effort, we developed a series of LM generation/filtering stages to create a larger version of the popular Winogender bias dataset. Our “Winogenerated” evaluation contains 50x as many examples as the original while obeying complex grammatical constraints.
-
Automated Generation of Yes-No Questions for LM Behavior Evaluation
By
–
We explored approaches with varying amounts of automation and human effort. In the simplest case, we generated thousands of yes-no questions for diverse behaviors just by instructing an LM (and filtering out bad examples with another LM).
-

Anthropic Explores Automating Language Model Evaluation
By
–
We explored approaches with varying amounts of automation and human effort. In the simplest case, we generated thousands of yes-no questions for diverse behaviors just by instructing an LM (and filtering out bad examples with another LM). Random examples of LM-written evals:
-

Automated Language Model Evaluations Using AI-Generated Tests
By
–
It’s hard work to make evaluations for language models (LMs). We’ve developed an automated way to generate evaluations with LMs, significantly reducing the effort involved. We test LMs using >150 LM-written evaluations, uncovering novel LM behaviors. https://
anthropic.com/model-written-
evals.pdf
… -

AI-Generating Algorithms: Alternative Path to Artificial General Intelligence
By
–
AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence Jeff Clune : https://
arxiv.org/abs/1905.10985 #AGI #AGIDebate #AIGA #ArtificialIntelligence -
Fine-tuning AI Models on Successful Response Patterns
By
–
Yep but wait till it's finetuned/trained on successful replies etc.
-
Licensing potential in datasets with minimal collection efforts
By
–
You're joking, but there's already 30M to 50M records with licenses (including Creative Commons) and they weren't seriously trying to collect license information. If they actually tried, they'd probably get to 10% 25% or more…
-

ChatGPT Generated Account Raises AI Authentication Concerns
By
–
Entire account ChatGPT generated replies, this is a first
