We're so early! even the current SoTA speech-to-speech models are so dumb compared to cascaded speech-to-speech (ASR + LLM + TTS) There's loads of low hanging fruits to pick still! Looking forward for open models to climb up this benchmark in 2025
Speech-to-Speech Models: Cascaded Approaches Still Outperform End-to-End
By
–
