Is there a good Llava-type vision model yet for video on @Replicate
? I'd like to use it for http://
TherapistAI.com because some people send videos of themselves I could hack it together myself:
– take the audio -> send to speech-to-text -> interpret it (like I do with audio
LLaVA video models on Replicate for AI therapy applications
By
–