Knowing the great people behind lmsys, I don't think there is anything really fishy in the gpt-4o-mini ranking like some people have been saying Likely more of a symptom that as models (and alignement methods) are getting much better it's been harder and harder to tell models
GPT-4o-mini Rankings and Model Alignment Methods Evaluation
By
–
