Top ML Papers of the Week (April 24 – 30): – AudioGPT
– Track Anything
– Agents Learn Soccer Skills
– Harnessing the Power of LLMs
– Scaling Transformer to 1M tokens
– A Cookbook of Self-Supervised Learning
…
MULTIMODAL AI
-
Top ML Papers Week: AudioGPT, Agents, Transformers, Self-Supervised Learning
By
–
-

ElevenLabs multilingual voice coming soon to WhatsApp
By
–
Your voice can now speak 8 languages. Elevenlabs voice cloning is now multilingual. This will likely be integrated into Whatsapp in a matter of weeks.
-
Life-like Avatars: Advanced 3D AI Technology Integration
By
–
Life-like Avatars, here we come pic.twitter.com/vnyHbjxbZm #Avatar #3D #MachineLearning #AI #ML #MixedReality #AR #VR #TechNews #IoT #technology #Programmers
— Catherine Adenle (@CatherineAdenle) 29 avril 2023Life-like Avatars, here we come #Avatar #3D #MachineLearning #AI #ML #MixedReality #AR #VR #TechNews #IoT #technology #Programmers
-

DeepFloyd Releases IF Text-to-Image Diffusion Model
By
–
Yesterday @deepfloydai released IF, a new text-to-image diffusion model that can (among other things) generate text inside images. We host the official demo on @huggingface Spaces Traffic has been massive, so the best time to try it is now, before the U.S. wakes up
-
Speech Bit Rate Surprisingly Low Even in Fast Speech
By
–
The bit rate of speech, even “fast” speech”, is absurdly low
-
Meta revolutionizes with realistic avatars
By
–
Meta jumps years ahead with lifelike Avatars.
— AI Breakfast (@AiBreakfast) 29 avril 2023
LLMs can be trained to speak like you.
Text-to-voice can sound like you.
The Simulation is coming. pic.twitter.com/dcU3KSMiGjMeta jumps years ahead with lifelike Avatars. LLMs can be trained to speak like you. Text-to-voice can sound like you. The Simulation is coming.
-
94 Million Embedded Passages Across 10 Languages
By
–
4/ Languages included: English, German, French, Spanish, Italian, Japanese, Arabic, Chinese (Simplified), Korean, and Hindi. That's a total of 94 million embedded passages!
-
Try out DeepFloyd IF model on Hugging Face Spaces
By
–
Try out the model here → https://
huggingface.co/spaces/DeepFlo
yd/IF
… -
Fully Parsed: AI-Generated Audio for Language Learning
By
–
@aeschylus_dev demoed Fully Parsed, which uses explorable audio generated by AI to teach people how to understand different languages.
— Replit ⠕ (@Replit) 28 avril 2023
He wanted to see how far he could get with audio processing on Replit, and after enabling Boosts, he had more than enough CPU to make it happen. pic.twitter.com/kLYPNyYWaV@aeschylus_dev demoed Fully Parsed, which uses explorable audio generated by AI to teach people how to understand different languages. He wanted to see how far he could get with audio processing on Replit, and after enabling Boosts, he had more than enough CPU to make it happen.
-

DeepFloyd AI Releases State-of-the-Art Text-to-Image Model
By
–
Announcing the release of DeepFloyd IF Our multimodal AI lab, @DeepFloydAI
, is publicly releasing their state-of-the-art text-to-image model. Learn more here → https://
stability.ai/blog/deepfloyd
-if-text-to-image-model
…