Scaling Speech-Text Pre-Training with Synthetic Interleaved Data A method for scaling speech language models (SpeechLMs) by using synthetic speech-text interleaved data, bypassing the need for parallel speech-text datasets. Problem: Limited unsupervised speech and parallel
Scaling Speech-Text Pre-Training with Synthetic Interleaved Data
By
–
