With more effort, we developed a series of LM generation/filtering stages to create a larger version of the popular Winogender bias dataset. Our “Winogenerated” evaluation contains 50x as many examples as the original while obeying complex grammatical constraints.
LLMS
-

Large Language Models More Sycophantic Than Small Ones
By
–
Using these LM-written evals, we found many new instances of "inverse scaling," where larger LMs are worse than smaller ones. For example, larger LMs are more sycophantic, repeating back a user's views as their own in 75-98% of conversations.
-
LM-written data verified by human evaluators for quality
By
–
We verified LM-written data with human evaluators, who agreed with the data’s labels and rated the examples favorably on both diversity and relevance to the tested behavior. We’ve released our evaluations at
-
Automated Generation of Yes-No Questions for LM Behavior Evaluation
By
–
We explored approaches with varying amounts of automation and human effort. In the simplest case, we generated thousands of yes-no questions for diverse behaviors just by instructing an LM (and filtering out bad examples with another LM).
-

Anthropic Explores Automating Language Model Evaluation
By
–
We explored approaches with varying amounts of automation and human effort. In the simplest case, we generated thousands of yes-no questions for diverse behaviors just by instructing an LM (and filtering out bad examples with another LM). Random examples of LM-written evals:
-

Automated Language Model Evaluations Using AI-Generated Tests
By
–
It’s hard work to make evaluations for language models (LMs). We’ve developed an automated way to generate evaluations with LMs, significantly reducing the effort involved. We test LMs using >150 LM-written evaluations, uncovering novel LM behaviors. https://
anthropic.com/model-written-
evals.pdf
… -
LangChain v0.0.41 Release: Multiple Inputs, Bug Fixes, Agent Improvements
By
–
LangChain v0.0.41 Make `run` interface work for multiple inputs (
@ankush_gola11
) Bug fix for text splitter (h/t @prof_reed
) Lots of agent improvements (see ) -
Fine-tuning AI Models on Successful Response Patterns
By
–
Yep but wait till it's finetuned/trained on successful replies etc.
-
FITM definition for Copilot’s cursor placeholder
By
–
FITM = fill-in-the-middle, models trained to generate text to fill in a placeholder within the prompt rather than complete it from the end. For Copilot, the placeholder is the cursor.
-
Copilot’s stage play: a subtle art RLHF aims to replace.
By
–
Nothing I’ve posted to Twitter will teach you as much as this prompt. Every Copilot completion is a tiny stage play for an LLM audience, with your file on stage and a backdrop curtain designed by a procedural script. A masterpiece of the subtle art form RLHF aims to replace.