AI Dynamics

Global AI News Aggregator

About

Critique of AI benchmark using AI evaluation on public questions

This was not a good benchmark before it was updated and it is not a good benchmark now. Having AIs evaluate the work of other AIs on publicly available questions from a different closed benchmark doesn’t tell you very much. And it is unclear how they establish the human ELO.

→ View original post on X — @emollick