Mega-TTS 2: Zero-Shot Text-to-Speech with Arbitrary Length Speech Prompts paper page: https://
huggingface.co/papers/2307.07
218
… Zero-shot text-to-speech aims at synthesizing voices with unseen speech prompts. Previous large-scale multispeaker TTS models have successfully achieved this goal with
MULTIMODAL AI
-

Mega-TTS 2: Zero-Shot Text-to-Speech with Arbitrary Length Prompts
By
–
-

Multimodal Transformer Foundation Model for Image and Text Generation
By
–
10/ Generative Pretraining in Multimodality – presents a new transformer-based multimodal foundation model to generate images and text in a multimodal context; enables performant multimodal assistants via instruction tuning.
-
AnimateDiff: Motion Module for Personalized Animated Image Generation
By
–
9/ AnimateDiff – appends a motion modeling module to a frozen text-to-image model, which is then trained and used to animate existing personalized models to produce diverse and personalized animated images. https://t.co/L4VOSdideP
— DAIR.AI (@dair_ai) 16 juillet 20239/ AnimateDiff – appends a motion modeling module to a frozen text-to-image model, which is then trained and used to animate existing personalized models to produce diverse and personalized animated images.
-

LLMs as General Pattern Machines: In-Context Learning and Robotics
By
–
6/ LLMs as General Pattern Machines – shows that without any additional training, LLMs are general sequence modelers, driven by in-context learning; applies 0-shot capabilities to robotics and transfer the pattern among words to actions.
-

NaViT: Flexible Vision Transformer for Any Resolution
By
–
5/ Patch n’ Pack: NaViT – a vision transformer for any aspect ratio and resolution through sequence packing; enables flexible model usage, improved training efficiency, and transfers to tasks involving image and video classification among others.
-
Zero-1-to-3: Single Image to 3D Object Generation
By
–
Zero-1-to-3: Zero-shot One Image to 3D Object
— AK (@_akhaliq) 16 juillet 2023
github: https://t.co/2QCZ1apGdv
demo: https://t.co/Z92zQae2Zn
learn to control the camera perspective in large-scale diffusion models, enabling zero-shot novel view synthesis and 3D reconstruction from a single image pic.twitter.com/dJSBD2qXMYZero-1-to-3: Zero-shot One Image to 3D Object github: https://
github.com/cvlab-columbia
/zero123
…
demo: https://
huggingface.co/spaces/cvlab/z
ero123-live
… learn to control the camera perspective in large-scale diffusion models, enabling zero-shot novel view synthesis and 3D reconstruction from a single image -

Guide for LLM-Powered Chart and Image Generation Systems
By
–
An excellent guide on working with LLM-powered chart and image generation systems!!
-

Computer Vision: Putting Humans at the Center of AI
By
–
3 Ways Computer Vision Will Put the Human in #AI in 2023 https://
bit.ly/3vy1UgE #ethics -

Computer Vision Embeddings for Machine Learning Applications
By
–
Computer Vision Embeddings for Machine Learning https://
bit.ly/3CzwZE7 #AI #MachineLearning #DeepLearning #LLMs #DataScience -

CM3Leon: Scaling Autoregressive Multi-Modal Models for Text and Images
By
–
Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning blog: https://
ai.meta.com/blog/generativ
e-ai-text-images-cm3leon/
… present CM3Leon (pronounced “Chameleon”), a retrieval-augmented, tokenbased, decoder-only multi-modal language model capable of generating and infilling both text and