Good question. I think the challenge here is that the evaluation is more challenging because of a) it's relative to all the other models
b) it's expensive to run the other models PS: Fun fact, I just wrote a tutorial on Elo ratings and constructing a leaderboard last week:
Challenges in Model Evaluation and Elo Rating Leaderboards
By
–