That makes sense. But you could also use a decoder-only architecture (with embedded image tokens as part of the input, as in LLaMA-Adapter, for example). (* This uses an encoder for the tokens, but it's still a decoder-only architecture due to the lack of cross-attention)
MULTIMODAL AI
-
Multi-epoch training effectiveness for Vision Transformers
By
–
A counter-argument is that multi-epoch training works pretty well for ViTs.
-
Transformers and LLMs: Beyond the Common Assumption
By
–
Not all transformers are LLMs, since transformers can also be used for computer vision. And not all LLMs are transformers as there are large language models based on – recurrent (RWKV, https://
arxiv.org/abs/2305.13048) – or convolutional (Hyena, https://
arxiv.org/abs/2302.10866) approaches. -

AI Primers and Guides on Diverse Machine Learning Topics
By
–
Really like these primers/guides on nearly any topic you can think of in AI. From model architectures, training techniques, speech, vision, NLP, multimodal, evaluation, etc… https://
aman.ai/primers/ai/ -
AudioCraft: Open-Source Music Generation Model
By
–
Composing music is as nearly hard as drawing/painting a good-looking photo/sketch for most people. We already have open frontier models for generating photo-realistic images but not so many models for music/audio generation.
— Jean de Dieu Nyandwi (@Jeande_d) 3 août 2023
AudioCraft is an open-source tool for generating music… https://t.co/DoNHeg62OrComposing music is as nearly hard as drawing/painting a good-looking photo/sketch for most people. We already have open frontier models for generating photo-realistic images but not so many models for music/audio generation. AudioCraft is an open-source tool for generating music
-

Brain2Music: Reconstructing Music from Human Brain Activity
By
–
Brain2Music: Reconstructing Music from Human Brain Activity https://
bit.ly/43D2Xuc https://
bit.ly/3Ovu1at -

Multimodal LLMs Enhance Medical Data Integration Capabilities
By
–
Medicine is inherently multimodal, involving a range of data such as medical images, clinical notes, lab tests, and more. Today on the Google Research blog, we discuss a spectrum of approaches for bringing multimodal capabilities to LLMs. Learn more → https://
goo.gle/3Ym4pzZ -
Language Models Guide Robots Performing Everyday Tasks
By
–
“Robot, set up table for pasta”. Check out our work using language models to guide robots performing everyday tasks. 👇 https://t.co/aQN130Wozv
— Fei-Fei Li (@drfeifei) 3 août 2023“Robot, set up table for pasta”. Check out our work using language models to guide robots performing everyday tasks.
-
Runway ML Launches Gen-2 Napping Dogs Footage Pack
By
–
Cozy up with our latest Gen-2 Footage Pack: Napping Dogs.
— Runway (@runwayml) 3 août 2023
Browse and download at https://t.co/pQCpPEb0Se pic.twitter.com/gkd7CDk8xzCozy up with our latest Gen-2 Footage Pack: Napping Dogs. Browse and download at http://
runwayml.com/footage-packs -

Meta-Transformer: Unified Multimodal Learning on GitHub
By
–
GitHub – invictus717/MetaTransformer: Meta-Transformer for Unified Multimodal Learning
https://bit.ly/3YaScOH
#AI #MachineLearning #DeepLearning #LLMs #DataScience