7/ Visual Instruction Tuning – uses language-only GPT-4 to generate multimodal language-image instruction-following data; applies instruction tuning and introduces LLaVA, a large multimodal model for general-purpose visual and language understanding.
LLaVA: Multimodal Model for Visual Language Understanding
By
–
