Perception is the bottleneck.
MLLMs can’t even transcribe a simple 5×5 grid from an image.
The best models got 0% accuracy on full-shape recognition.
They see… but don’t understand.
MLLMs fail at simple grid transcription and shape recognition tasks
By
–
