Learn how we turned a Vision Transformer image encoder into an efficient video backbone using sparse video tubes (3D grid-based cuboids with learnable visual representations of video samples), reducing compute needs and achieving competitive results → https://
goo.gle/435CTYQ
Vision Transformer Efficient Video Backbone Using Sparse Tubes
By
–
