
On M-Portal: 12 frontier models were tested.
Every model scored near random chance.
Even GPT-4o got only 4.1% on a key task.
GPT-o3 was best, with a mere 17.6% on the easy version. This is blindness.
By
–


On M-Portal: 12 frontier models were tested.
Every model scored near random chance.
Even GPT-4o got only 4.1% on a key task.
GPT-o3 was best, with a mere 17.6% on the easy version. This is blindness.