AI Dynamics

Global AI News Aggregator

About

RLHF Annotation Bias and Model Vocabulary Development

I don't know that RLHF would bias that kind of thing – my mental model is that annotators are shown two answers to the same prompt and asked which is "best", so if none of the test prompts happened to touch on the concept of a roadside kiosk that vocabulary wouldn't be affected

→ View original post on X — @simonw