AI Dynamics

Global AI News Aggregator

About

LLaVA: Multimodal Model for Visual Language Understanding

7/ Visual Instruction Tuning – uses language-only GPT-4 to generate multimodal language-image instruction-following data; applies instruction tuning and introduces LLaVA, a large multimodal model for general-purpose visual and language understanding.

→ View original post on X — @dair_ai