AI Dynamics

Global AI News Aggregator

About

Humanity’s Last Exam: No AI model scores above 10%

Examples from Humanity’s Last Exam, designed by CAIS & Scale to be the *final* broad-subject, closed-end Q&A exam for LLMs written by humans. Today, no model scores above 10%. This is how hard a test has to be for frontier AI to score that low — in Classics, Ecology, and C.S.:

→ View original post on X — @goodside