In my experiment, a preference predictor was able to pick up the performance patterns of different models. One pattern is that for simple prompts, weak models can do (nearly) as well as strong models. For more challenging prompts, however, users are much more likely to prefer
Weak Models Match Strong Models on Simple Prompts
By
–
