Having played all of these games, I feel strongly that I would have scored >95% in a real testing session. Even the #1 human tester replay tends to be very, very far from optimal, and we're using #2 as baseline, so it's easy to score 100% on an given environment. You don't need
@fchollet
-
AGI Must Build Its Own Harness for True Generality
By
–
AGI will make its own harness (or whatever else it needs to solve a new problem). As long as you need a human engineer to handcraft a task-specific harness/system for each new problem, AI isn't general. It's an automation tool to be wielded by software engineers. Harness-related
-
New AGI Eval Focuses Research Efforts on Critical Gaps
By
–
If you care about the rate of AGI progress, you should be excited about a new eval that focuses research efforts by pointing out important gaps & providing a way to measure progress towards fixing them If instead you only care about having your preconceptions confirmed, too bad
-
ARC-AGI-3 Environments Mirror Scientific Method for Breakthrough AI
By
–
Many people expect that current AI is ready to cure cancer and do breakthrough new science. ARC-AGI-3 envs are like a microcosm of the scientific method: you must observe a tiny world, form a theory of how it works, test it, iterate until correct. Over the course of a few
-
Open-sourcing AI testing dataset with full transparency
By
–
We have published an extensive technical report on our methodology and we will be open-sourcing the full human testing dataset. We have always be maximally transparent about our process and our reasoning.
-
AI Systems Fall Short of Human Job Performance Standards
By
–
Virtually every human job on earth has a higher bar. These are not very high expectations for AI systems that claim to be able to do everything humans can.
-
Defining ASI: Super Intelligence Beyond Human Performance
By
–
"2+ people can do it out of an unfiltered pool of 10 people that might well be a below-average sample" is not the sign of a insurmountable challenge. It's not certainly where I would set the bar for "super intelligence". ASI is when AI is better than *every single human* — for
-
ARC-AGI Benchmark Standards and Human Performance Expectations
By
–
This is a very low bar, objectively. The claim is obviously not that 100% of humans could solve 100% of the games — that would be silly, and it wouldn't be true either of ARC-AGI1 or 2, nor of any AI benchmark that has ever been used in the field. Not even MNIST can be 100%
-
ARC-AGI-3 Environments Meet Human Feasibility Standards
By
–
To be clear, all ARC-AGI-3 environments are feasible by humans with no prior ARC-AGI-3-specific training. Our bar for feasibility is the following… Each environment was seen by 10 human testers. If 2 testers could independently clear it (successfully solving *all* levels in
-
ARC-AGI-4 Benchmark Release Scheduled Early 2027
By
–
For those wondering about ARC-AGI-4 timing: it will be released in early 2027. We are aiming for a yearly release schedule for new benchmarks. We are also aiming for each new benchmark to be fully unsaturated upon release, and to target the most important unanswered research