It generates a series of codes based on the phonemes and an audio recording of the speaker's voice.
— AI Breakfast (@AiBreakfast) 7 janvier 2023
These codes are then turned directly into a waveform, which allows it to generate speech that sounds like the speaker from just a small amount of reference audio.
Example (🔊): pic.twitter.com/HQvjJUnbiH
It generates a series of codes based on the phonemes and an audio recording of the speaker's voice. These codes are then turned directly into a waveform, which allows it to generate speech that sounds like the speaker from just a small amount of reference audio. Example ():