Introducing Vid2Seq, a visual language model for dense video captioning that simply predicts all event boundaries and captions as a single sequence of tokens. Learn more about how it achieves state-of-the-art results on various benchmarks → https://t.co/CgQXVBnNYs pic.twitter.com/oQ1fXwEBAj
— Google AI (@GoogleAI) 17 mars 2023
Introducing Vid2Seq, a visual language model for dense video captioning that simply predicts all event boundaries and captions as a single sequence of tokens. Learn more about how it achieves state-of-the-art results on various benchmarks → https://
goo.gle/3JNz97v