10/ MultiModal-GPT – a vision and language model for multi-round dialogue with humans; the model is fine-tuned from OpenFlamingo, with LoRA added in the cross-attention and self-attention parts of the language model.
MultiModal-GPT: Vision Language Model for Multi-Round Dialogue
By
–