AI Dynamics

Global AI News Aggregator

About

AutoGaze: Efficient Video Understanding with Selective Attention

Humans can see in high-res, high-FPS in real-time. Why can't VLMs? Introducing AutoGaze: ViTs/VLMs "gaze" only at key video regions! Up to 4-100x token savings, 19x speedup, and enables scaling to 4K-res 1K-frame videos. 📄 arxiv.org/abs/2603.12254 🌐 autogaze.github.io 🤗 huggingface.co/collections/b… (1/n)🧵

→ View original post on X — @berkeley_ai, 2026-03-24 18:46 UTC