What if a single model could generate audio from video, text, or both — with no trade-offs? Researchers from Tsinghua University, Monash University, and Shengshu AI present Omni2Sound. They built SoundAtlas (470k high-alignment pairs) and a three-stage training schedule to
Omni2Sound: single model for audio from video and text
By
–
