Unlike standard Q&A style benchmarks that eventually saturate, these tests auto get harder as the models get better. Great to have these verifiable ways to measure progress toward AGI. Goal is to add 100s of games covering many aspects of intelligence, with an overall leaderboard