


Examples from Humanity’s Last Exam, designed by CAIS & Scale to be the *final* broad-subject, closed-end Q&A exam for LLMs written by humans. Today, no model scores above 10%. This is how hard a test has to be for frontier AI to score that low — in Classics, Ecology, and C.S.:
