In early 2022 @dribnet
's Pixray was the first text-to-image model on Replicate to reach thousands of runs. Today it's been run a total of 1.3mil times. @dribnet was actually the first Replicate user to request that we build an API instead of just a web ui. The rest is history
MULTIMODAL AI
-

Pixray reaches 1.3 million runs on Replicate platform
By
–
-

DALL-E Mini Launch: July 2021 Breakthrough by Boris Dayma
By
–
Fast-forward a few months to July 2021: @borisdayma published DALL·E Mini https://
x.com/borisdayma/sta
tus/1421117516605267968
… This is an excellent breakdown of how it was all put together: https://
wandb.ai/dalle-mini/dal
le-mini/reports/DALL-E-Mini-Explained–Vmlldzo4NjIxODA
… -

OpenAI’s DALL·E 2 and the Rise of Diffusion Models
By
–
In April 2021, OpenAI announced DALL·E 2 https://
openai.com/dall-e-2 and detailed their transition toward using diffusion models. It felt like the dream was coming true! You could prompt it for a photo of a cat wearing a red hat, and get a cat wearing a red hat! -

BigSleep Colab Notebook Shared by advadnoun
By
–
Then just a week later @advadnoun shared another Colab notebook named The BigSleep.
-

DeepDaze: Text-to-Image Generation via Colab
By
–
A few weeks after the release of the OpenAI papers, @advadnoun shared DeepDaze, a Colab notebook that generated images from text. Originally shared here:
https://x.com/advadnoun/status/1348375026697834496
… The images were quite abstract, but legible enough to inspire more experimentation. -

TokenSplit: Discrete Token Speech Separation Model at Interspeech
By
–
At 3:30pm today, drop by the #Interspeech2023 Google booth to hear @ScottTWisdom & @HakanErdoganPhD discuss TokenSplit, a speech separation model that uses discrete token sequences representing audio signals & can improve output by using transcripts or separated audio inputs.
-
AI-Powered AR Portals Enable Immersive Remote Experiences
By
–
Estos portales son una experiencia de realidad aumentada en tiempo real impulsadas por ia y pueden llevarte a cualquier lugar que imagines.
— Juan Merodio (@juanmerodio) 22 août 2023
Estos portales nos permitirán entrar en lugares remotos sin movernos del sitio creando la ilusión de haberlos visitado físicamente.… pic.twitter.com/scrcMHUYdOEstos portales son una experiencia de realidad aumentada en tiempo real impulsadas por ia y pueden llevarte a cualquier lugar que imagines. Estos portales nos permitirán entrar en lugares remotos sin movernos del sitio creando la ilusión de haberlos visitado físicamente.
-
AudioCraft: Unified Audio Modeling Solution for AI
By
–
AudioCraft: A simple one-stop shop for audio modeling https://
bit.ly/3Kt8Bby #AI #MachineLearning #DeepLearning #LLMs #DataScience -
EXPRESSO: Benchmark for Discrete Expressive Speech Resynthesis
By
–
EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis https://
bit.ly/3si2sc6 A high-quality expressive speech dataset for textless speech synthesis — includes read speech & improvised dialogues rendered in 26 spontaneous expressive styles. -
Robust End-to-End Spoken Language Understanding with Modality Confidence
By
–
Modality Confidence Aware Training for Robust End-to-End Spoken Language Understanding https://
bit.ly/45cZLa5 A novel E2E SLU system that enhances robustness to ASR errors by fusing audio + text representations based on est. modality confidence of ASR hypotheses.