AI Dynamics

Global AI News Aggregator

About

Vision Transformer Efficient Video Backbone Using Sparse Tubes

Learn how we turned a Vision Transformer image encoder into an efficient video backbone using sparse video tubes (3D grid-based cuboids with learnable visual representations of video samples), reducing compute needs and achieving competitive results → https://
goo.gle/435CTYQ

→ View original post on X — @googleai