Por cierto el modelo, como buen modelo multimodal, es capaz de condicionar su resultado al input de audio. En este caso este vídeo no tenía ningún prompt de input mas que el vídeo y su audio. pic.twitter.com/G4IPOexVXS
— Carlos Santana (@DotCSV) 19 mai 2026
By the way, as a robust multimodal model, it can condition its output on audio input. In this instance, the video had no input prompt other than the video itself and its audio.