Audio should be easier. You could for example only transcribe voices that match a voice that has given permission in that voice. You’d still need AI to analyze whether it was AI generated but I suspect doing that on voice is easier than text.
Voice Authentication and AI-Generated Audio Detection Methods
By
–