MARBLE exposes the gap between mimicking intelligence and real cognition. Planning, perception, execution is still human. Scale won't solve it. Solve MARBLE, and you build the next trillion-dollar model. link to paper: https://
arxiv.org/pdf/2506.22992
AGI
-
MARBLE exposes the gap between mimicking intelligence and real cognition
By
–
-

AI models can’t handle multi-step search in CUBE task
By
–
Reasoning is worse.
The CUBE task has 188 million possible solutions.
Even when perception is bypassed, models collapse under combinatorics.
They cannot hold multi-step chains.
They cannot search. -

12 frontier models fail key tasks; GPT-4o only 4.1%, GPT-o3 17.6%
By
–

On M-Portal: 12 frontier models were tested.
Every model scored near random chance.
Even GPT-4o got only 4.1% on a key task.
GPT-o3 was best, with a mere 17.6% on the easy version. This is blindness. -

MARBLE benchmark demands spatial reasoning; even o3 fails
By
–
MARBLE isn’t trivia. It’s not "what's in this picture?" It’s: "Given this environment, how do you escape the room… step-by-step?"
Or
"Can you assemble a cube from 6 jigsaw pieces under spatial constraints?" Answer: no.
Not even o3 can. -
O3 AI Model Achieves Impressive Progress Milestone
By
–
(None of that is to detract from how impressive it is that O3 has gotten as far as it has.)
-
AI Sycophancy Problem: Need for Honest Critical Feedback
By
–
A lot of people assume that there are more objective answers out there than there actually are. I don't think sycophancy is a huge issue in math, but you want advice? Feedback on writing? Help with an idea or project? Lots of areas where current mode default to too nice.
-
Semantic Understanding in AI Evaluation Metrics
By
–
Our family rules are that it only counts as "solved" if you figure out the final category descriptor correctly (i.e semantically close, not identical wording). NYT's Connections Bot mentions the same thing sometimes. I don't think the current evals check this AFAIK.
-
Symbol Manipulation Critical for AI Progress Beyond Current Limitations
By
–
literally my critique (1998-2025) was that if you don’t add symbol-manipulation you fail in certain ways, and the systems failed in those ways. once you add in symbol-manipulation, things get more interesting, and the rate-limiting step becomes how you leverage those new tools.
-
AGI Alpha: Ascending Together with Artificial General Intelligence
By
–
[ α‑AGI Ascension ] “We Choose to Ascend with AGI” https://
agialphaagent.com Together, we choose not merely to witness the dawn of #AGI, but to ascend with it. #AGIALPHA #AscendWithAGI
