AI Dynamics

Global AI News Aggregator

About

Scale AI and CAIS Release EnigmaEval Challenging Reasoning Benchmark

On the heels of Humanity's Last Exam, @scale_AI & @cais have released a new very-hard reasoning eval: EnigmaEval: 1,184 multimodal puzzles so hard they take groups of humans many hours to days to solve. All top models score 0 on the Hard set, and <10% on the Normal set

→ View original post on X — @alexandr_wang