Generative Pretraining in Multimodality paper page: https://
huggingface.co/papers/2307.05
222
… present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-modality or multimodal data
Emu: Multimodal Foundation Model for Image and Text Generation
By
–
