"Intermediate representation" refers to a representation of the music generated from the text prompt. The generator model first turns the text prompt into this intermediate representation, which captures elements of the music such as: -Genre
-Tempo
-Instruments
-Mood
-Era
MULTIMODAL AI
-
Intermediate representation in music generation from text prompts
By
–
-
Zero-shot classification matches text to music without training
By
–
The model knows how to connect text and music to match each of these sentences with a specific music clip, without having to specifically train the model for each clip.
— AI Breakfast (@AiBreakfast) 9 février 2023
This process is called zero-shot classification. pic.twitter.com/KWerrsnMWjThe model knows how to connect text and music to match each of these sentences with a specific music clip, without having to specifically train the model for each clip. This process is called zero-shot classification.
-
Future of music creation via text prompts using diffusion models
By
–
In the not-too-distant future, anyone will be able to create any type of music they want via text prompts.
— AI Breakfast (@AiBreakfast) 9 février 2023
Here's a look into recent research on diffusion models for generating high quality music audio from text prompts:
(more examples below) pic.twitter.com/ywgJRVoxLxIn the not-too-distant future, anyone will be able to create any type of music they want via text prompts. Here's a look into recent research on diffusion models for generating high quality music audio from text prompts: (more examples below)
-
Training generator and cascader models on 300k+ hours of audio
By
–
The 300k+ hours of audio clips were used to train a "generator model" that turns the text into an intermediate representation, and a "cascader model" that uses this intermediate representation to produce high-quality audio.
-

Noice2Music generates 30-second music from text inputs
By
–
"Noice2Music" is a project from Google Research that generates a short 30 second piece of music based on text inputs.
-

Flair: Quick Branded Content Design with AI
By
–
3. Quick Branded Content Design with Flair Create eye-catching content in no time with Flair. https://
flair.ai -
Computer Vision and Robotics Integration
By
–
Computer Vision + Robotics. This is going to be a great one
-
Generative AI transforms maps into realistic 3D with NeRF
By
–
L’IA generative permet aussi de transformer la manière dont on interagis avec Map. Par exemple en créant des cartes du mondes en 3D et “réel” et le pire c’est l’intérieur des lieux 🤯🤯 regardez la vidéo du resto. C’est pas une “vrai” vidéo omg ça utilise du nerf. pic.twitter.com/s7m5WokLav
— Defend Intelligence (Anis Ayari) (@DFintelligence) 8 février 2023Generative AI also allows transforming the way we interact with Maps. For example, by creating 3D and 'real' maps of the world, and the craziest part is the interior of places – watch the restaurant video. It's not a 'real' video, omg it uses NeRF.
-

Understanding all information with transformers according to Google
By
–

The goal is clearly to better understand all available information (images, texts, etc.). Google reminds that it is thanks to transformers.
-
New Audio Speech Model Announcement Expected Today
By
–
A little bird told me that @mhollemans might be announcing a very cool audio/ speech model today! Stay tuned..