
Good post with maybe big implications for interpreting o3’s ARC-AGI scores: TLDR:
– LLMs bad at ARC because they can’t perceive large text grids
– Solve rates fall as task size in pixels rises, but much later for o3
– 80% of pass@1 o1-mini tasks fail when grids enlarged 2x

