Through 1. Vision encoders that are not LLMs. They are actually Joint Embedding Architectures that embed images and text description in the same space
2. Painfully exhaustive training on enormous amounts of declarative facts about the physical world. You can train them to answer
Joint Embedding Architectures: Vision Encoders Beyond LLMs
By
–