AI Dynamics

Global AI News Aggregator

About

Humanity’s Last Exam Benchmark Flawed AI Measurement Concerns

This is bad for AI measurement. As other AI benchmarks have become saturated, model makers have turned to Humanity’s Last Exam as a good measure of AI ability. Except a careful review suggests many of the exam questions have incorrect “right” answers. Benchmarking is hard.

→ View original post on X — @emollick