5/ Patch n’ Pack: NaViT – a vision transformer for any aspect ratio and resolution through sequence packing; enables flexible model usage, improved training efficiency, and transfers to tasks involving image and video classification among others.
NaViT: Flexible Vision Transformer for Any Resolution
By
–
