MARBLE exposes the gap between mimicking intelligence and real cognition. Planning, perception, execution is still human. Scale won't solve it. Solve MARBLE, and you build the next trillion-dollar model. link to paper: https://
arxiv.org/pdf/2506.22992
LLMS
-
MARBLE exposes the gap between mimicking intelligence and real cognition
By
–
-

AI models can’t handle multi-step search in CUBE task
By
–
Reasoning is worse.
The CUBE task has 188 million possible solutions.
Even when perception is bypassed, models collapse under combinatorics.
They cannot hold multi-step chains.
They cannot search. -

MLLMs fail at simple grid transcription and shape recognition tasks
By
–
Perception is the bottleneck.
MLLMs can’t even transcribe a simple 5×5 grid from an image.
The best models got 0% accuracy on full-shape recognition.
They see… but don’t understand. -

All models failed M-Cube; GPT-o3 barely 72% after huge reduction
By
–

On M-Cube: All models failed.
Zero percent on full tasks.
Even with 10,000+ tokens of “reasoning.” On the simplified version? GPT-o3 barely crossed 72% – after reducing the search space by 5 million-fold. -

12 frontier models fail key tasks; GPT-4o only 4.1%, GPT-o3 17.6%
By
–

On M-Portal: 12 frontier models were tested.
Every model scored near random chance.
Even GPT-4o got only 4.1% on a key task.
GPT-o3 was best, with a mere 17.6% on the easy version. This is blindness. -

M-Portal and M-Cube: MLLMs fail at visual, planning, spatial logic
By
–

There are two tasks: M-Portal: Reason like a Portal 2 player. M-Cube: Rebuild a 3D object from 2D slices. Both require: Visual perception, Multistep planning, and Spatial logic MLLMs today fail at all three.
-

AI Reasoning Limits: MARBLE Benchmark Highlights Multimodal LLM Weaknesses
By
–
Multimodal LLMs can write essays.
They can chat, caption, and even summarize papers. But give them a Portal map or a 3D puzzle, they break instantly. Zero percent accuracy. Welcome to the edge of AI reasoning: The MARBLE Benchmark -

MARBLE benchmark demands spatial reasoning; even o3 fails
By
–
MARBLE isn’t trivia. It’s not "what's in this picture?" It’s: "Given this environment, how do you escape the room… step-by-step?"
Or
"Can you assemble a cube from 6 jigsaw pieces under spatial constraints?" Answer: no.
Not even o3 can. -
Grok 3 and 4 System Prompt Transparency Concerns
By
–
Given the many issues with the system prompt, I really want to see the current version for Grok 3 (X answerbot) and Grok 4 (when it comes out). Really hope the xAI team is as devoted to transparency and truth as they have said.
-
Distillation increases model capacity efficiency
By
–
it seems pretty likely distillation would increase capacity if that’s what u mean