The language model was trained on a huge amount of data (60,000 hours of English speech) and combines it with just 3 seconds of a person's voice and uses that to synthesize new, high-quality speech that sounds like the original speaker. pic.twitter.com/tIaimhfQ9R
— AI Breakfast (@AiBreakfast) 7 janvier 2023
The language model was trained on a huge amount of data (60,000 hours of English speech) and combines it with just 3 seconds of a person's voice and uses that to synthesize new, high-quality speech that sounds like the original speaker.