Frontier LLM have superhuman text-based world knowledge. Frontier image / video models have superhuman vision-based world knowledge (e.g. Genie). But current frontier VLMs are still absolutely clown shoes. Why? Relative scarcity of image:text pairs (while there is plenty of text
Why Frontier VLMs Lag Behind LLMs and Vision Models
By
–