AI Dynamics

Global AI News Aggregator

About

MultiModal-GPT: Vision Language Model for Multi-Round Dialogue

10/ MultiModal-GPT – a vision and language model for multi-round dialogue with humans; the model is fine-tuned from OpenFlamingo, with LoRA added in the cross-attention and self-attention parts of the language model.

→ View original post on X — @dair_ai