If you’re using out of the box LLMs as a reward model/sample rater, here’s a trick to get way better performance: Don’t ask the LLM to rate examples on a numerical scale (i.e. 1-5). The model will almost always choose 1 or 5. Instead, use words as rating options (“very bad”,
Better LLM Reward Model Performance Without Numerical Scales
By
–