What if standard vision-language models already understand 3D—without complex architecture changes or special losses? Meta and Princeton University present VLM³, showing that VLMs are native 3D learners. Their recipe: unify camera focal lengths, use text-based pixel references,
VLM³ proves vision-language models are native 3D learners
By
–
