3/ Researchers ran this test on the top AI models. GPT-5, Claude Opus 4.1, Gemini 2.5, and others. On short lists they aced it. 90% and up. Then the researchers made the lists longer.
AI models ace short lists but struggle with longer ones
By
–
By
–
3/ Researchers ran this test on the top AI models. GPT-5, Claude Opus 4.1, Gemini 2.5, and others. On short lists they aced it. 90% and up. Then the researchers made the lists longer.