View the upgrade in definition between SDXL beta (left) and SDXL 0.9 in these examples: Greater detail with human hands
MULTIMODAL AI
-

Stability AI Releases SDXL 0.9 with Major Improvements
By
–
Introducing the latest release from Stability AI: Breaking barriers with #SDXL 0.9! SDXL 0.9 produces massively improved text-to-image and composition detail over the beta release and provides a leap in use cases for generative AI imagery. #StabilityAI Unleash your creativity
-
StyleSync: AI Project for Creative Content Generation
By
–
You can find more details on the project page. https://
hangz-nju-cuhk.github.io/projects/Style
Sync
… -

Awesome Multimodal Large Language Models: Papers and Datasets
By
–
Awesome Multimodal Large Language Models A curated list of the latest papers and datasets in multimodal large models. Provides an excellent listing of multimodal papers, codes, and demos. Papers about topics like instruction tuning, in-context learning, chain-of-thought(CoT),
-
StyleSync: High-Fidelity Generalized Personalized Lip Sync
By
–
Demo video for the paper: StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator (CVPR 2023).
-

Google Reveals REVEAL: Retrieval-Augmented Visual-Language Model
By
–
At #CVPR2023? Drop by the Google booth at 12:30pm today to hear @alirezafathi talk about REVEAL (https://t.co/MTP69suQ6d), an end-to-end retrieval-augmented visual-language model that learns to use multi-source, multi-modal data to answer knowledge-intensive queries. pic.twitter.com/XqunzICYya
— Google AI (@GoogleAI) 21 juin 2023At #CVPR2023? Drop by the Google booth at 12:30pm today to hear @alirezafathi talk about REVEAL (
http://
goo.gle/3qcZwwc), an end-to-end retrieval-augmented visual-language model that learns to use multi-source, multi-modal data to answer knowledge-intensive queries. -
MAGVIT: Single Transformer for Ten Video Generation Tasks
By
–
Introducing MAGVIT, a single transformer designed to address ten video generation tasks, including future image animation, video editing and video outpainting. Stop by the #CVPR2023 Google booth at 10am today to see it in action! More details at: https://t.co/nd7Gew2ozv pic.twitter.com/8J35j5fnIG
— Google AI (@GoogleAI) 21 juin 2023Introducing MAGVIT, a single transformer designed to address ten video generation tasks, including future image animation, video editing and video outpainting. Stop by the #CVPR2023 Google booth at 10am today to see it in action! More details at: https://
magvit.cs.cmu.edu -
Facebook Research MMS Examples on GitHub Repository
By
–
github: https://
github.com/facebookresear
ch/fairseq/tree/main/examples/mms
… -
Facebook AI Multilingual Speech Recognition Model
By
–
blog: https://ai.facebook.com/blog/multilingual-model-speech-recognition/
-
Documentation of MMS Transformers on Hugging Face
By
–
docs: https://huggingface.co/docs/transformers/main/en/model_doc/mms
