Google presents Zero-shot Cross-lingual Voice Transfer for TTS discuss: https://
huggingface.co/papers/2409.13
910
… In this paper, we introduce a zero-shot Voice Transfer (VT) module that can be seamlessly integrated into a multi-lingual Text-to-speech (TTS) system to transfer an individual's
GENERATIVE AI
-

Google Zero-shot Cross-lingual Voice Transfer for TTS
By
–
-

Nvidia MaskedMimic: Unified Physics-Based Character Control
By
–
Nvidia presents MaskedMimic
— AK (@_akhaliq) 24 septembre 2024
Unified Physics-Based Character Control Through Masked Motion Inpainting
discuss: https://t.co/Q2Tx01NP2I
Crafting a single, versatile physics-based controller that can breathe life into interactive characters across a wide spectrum of scenarios… pic.twitter.com/E3lADK3FcZNvidia presents MaskedMimic Unified Physics-Based Character Control Through Masked Motion Inpainting discuss: https://
huggingface.co/papers/2409.14
393
… Crafting a single, versatile physics-based controller that can breathe life into interactive characters across a wide spectrum of scenarios -

Self-Supervised Audio-Visual Soundscape Stylization Techniques
By
–
Self-Supervised Audio-Visual Soundscape Stylization
— AK (@_akhaliq) 24 septembre 2024
discuss: https://t.co/mEqithUXlD
Speech sounds convey a great deal of information about the scenes, resulting in a variety of effects ranging from reverberation to additional ambient sounds. In this paper, we manipulate input… pic.twitter.com/151CoB2BaMSelf-Supervised Audio-Visual Soundscape Stylization discuss: https://
huggingface.co/papers/2409.14
340
… Speech sounds convey a great deal of information about the scenes, resulting in a variety of effects ranging from reverberation to additional ambient sounds. In this paper, we manipulate input -
Meta Connect 2024: Metaverse and Generative AI Event
By
–
Meta Connect 2024: How to watch the #metaverse and generative AI event https://
techcrunch.com/2024/09/23/met
a-connect-2024-how-to-watch-the-metaverse-and-generative-ai-event/
… via @techcrunch -

SolarPro: Most Powerful LLM on Single GPU
By
–
@upstageai #SolarPro has been featured in the most popular tldr newsletter among developers and startups (600,000 viewers). The article is titled "Most powerful LLM on a single GPU (12 minute read)". Please give it a read! https://
a.tldrnewsletter.com/web-version?ep
=1&lc=3bf8d792-bd90-11ee-817b-b777f755fa8b&p=703435f8-799a-11ef-8b11-371f32244928&pt=campaign&t=1727097241&s=dcda0944a0dad5602d8b27308915cf426959da99e53da7f4f12b2e6ffb0fd41f
… -

New Gemini models expected to launch tomorrow
By
–

Looks like new Gemini models are dropping tomorrow and not Opus 3.5
-

User experience with AI memory and personalization in voice mode
By
–
Personalisation isn't a new feature, but immediately after asking "How is it going?", I got a question back about one of the events I've saved into memory in the past In voice mode, it feels very different (even if it is just standard).
-

Testing new voice features in ChatGPT read aloud
By
–
The Voice selector is still very raw on desktop and these extra voices are not yet available. I've added some extra values for demo proposes. More info https://
testingcatalog.com/chatgpt-hints-
at-up-to-8-new-voices-available-for-read-aloud/
… -

Turbo LoRA: Fast Inference with Customized Models
By
–
Last chance to save your spot Learn how to get the best of both worlds: super fast #inference + high quality models customized for your use case! Join us tomorrow, to learn about Turbo LoRA, a new approach to fine-tuning that increases model #throughput by 2-3x, lowers
-
Evolution of Deep Learning: From Visualization to GANs and VAEs
By
–
Deepdream and neural style transfer are from around the same time period and used deep neural nets. The precursor to them was computer vision stuff on visualization like https://
arxiv.org/pdf/1112.6209 & https://
arxiv.org/abs/1312.6034 GANs were also predated by VAEs. VAEs might have been the