I'm curious about audio models like this — I know there's a lot of excitement about them, but I don't get it yet. Do you find in practice it frequently gives you benefits over doing speech recognition and a text model?
Audio Models vs Speech Recognition: Practical Benefits Analysis
By
–