AI Dynamics

Global AI News Aggregator

About

Joint Embedding Architectures: Vision Encoders Beyond LLMs

Through 1. Vision encoders that are not LLMs. They are actually Joint Embedding Architectures that embed images and text description in the same space
2. Painfully exhaustive training on enormous amounts of declarative facts about the physical world. You can train them to answer

→ View original post on X — @ylecun