What if an AI could generate a video with perfectly synced speech, not just background noise? Researchers from The University of Hong Kong & ByteDance present JoVA. It uses a single, streamlined transformer where video and audio data directly interact in every layer, avoiding
JoVA: AI Generates Perfectly Synced Video Speech
By
–
