Instead of turning text into a series of specific sounds (called phonemes) and then into a visual representation of the sound called a mel-spectrogram, and then into a waveform (a digital representation of sound that can be played through a speaker) VALL-E takes a shortcut:
Technical architecture of the VALL-E generative audio model
By
–