EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis https://
bit.ly/3si2sc6 A high-quality expressive speech dataset for textless speech synthesis — includes read speech & improvised dialogues rendered in 26 spontaneous expressive styles.
@aiatmeta
-
EXPRESSO: Benchmark for Discrete Expressive Speech Resynthesis
By
–
-
Robust End-to-End Spoken Language Understanding with Modality Confidence
By
–
Modality Confidence Aware Training for Robust End-to-End Spoken Language Understanding https://
bit.ly/45cZLa5 A novel E2E SLU system that enhances robustness to ASR errors by fusing audio + text representations based on est. modality confidence of ASR hypotheses. -

ImageBind: Meta AI’s Multimodal Learning Across Six Modalities
By
–
ImageBind by Meta AI is capable of binding information from six modalities. It equips machines with a holistic understanding that connects objects in photos with how they will sound, their 3D shape, temperature & how they move.
— AI at Meta (@AIatMeta) 16 août 2023
More on this research ➡️ https://t.co/yzoIrLRu2p pic.twitter.com/nHOxevSA51ImageBind by Meta AI is capable of binding information from six modalities. It equips machines with a holistic understanding that connects objects in photos with how they will sound, their 3D shape, temperature & how they move. More on this research https://
bit.ly/46hAJaY -

Meta AI Introduces Shepherd: A Model for Critiquing and Refining Responses
By
–
New research from Meta AI — Shepherd is a language model specifically tuned to critique model responses & suggest refinements. It goes beyond the capabilities of untuned models to identify diverse errors & suggest improvements. Read the paper https://
bit.ly/3QC4AW3 -
New Demo Released with Additional Details Available
By
–
New demo available now at the link in Anwar's tweet. More details on this work ➡️ https://t.co/eg117vazQo https://t.co/cDRopcEy1O
— AI at Meta (@AIatMeta) 14 août 2023New demo available now at the link in Anwar's tweet. More details on this work https://
bit.ly/45r94mD -
Anyscale Team Shares Guide on Fine-Tuning Llama 2
By
–
Great read from the team at @anyscalecompute on fine-tuning Llama 2.
-
Meta AI Launches AudioCraft: Generative Audio Models Suite
By
–
AudioCraft by Meta AI is a one-stop codebase for generative audio consisting of three models: MusicGen, AudioGen & EnCodec — supporting both compression + generation of high-quality music and sound effects from text.. Get the code
-
Llama 2 Now Available on IBM watsonx.ai Platform
By
–
We’re excited for even more people to be able to build with Llama 2 on https://t.co/ZEoYo4mPNn! https://t.co/RruqN3rTzC
— AI at Meta (@AIatMeta) 9 août 2023We’re excited for even more people to be able to build with Llama 2 on http://
watsonx.ai! -
PUG Datasets Improve CV Model Robustness Assessment
By
–
Even the most recent CV models are still struggling w/ basic object classification tasks that are easy for humans. With the new PUG datasets, we can correctly assess and find ways to improve the robustness of these models. Research Paper
-
Meta AI PUG: Photorealistic Controllable Datasets for Model Evaluation
By
–
Today we're sharing our work on PUG, new research from Meta AI on photorealistic, semantically controllable datasets using Unreal Engine for robust model evaluation.
— AI at Meta (@AIatMeta) 9 août 2023
More details & dataset downloads ➡️ https://t.co/rpBBmhNyyK pic.twitter.com/avZEQAysXhToday we're sharing our work on PUG, new research from Meta AI on photorealistic, semantically controllable datasets using Unreal Engine for robust model evaluation. More details & dataset downloads https://
bit.ly/45na9M6