AI Dynamics

Global AI News Aggregator

About

LLaVA-ST: Multimodal LLM for Spatial-Temporal Video Understanding

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding LLaVA-ST is a multimodal large language model designed for spatial-temporal understanding in videos, addressing challenges in coordinate alignment and feature compression. Problem:

→ View original post on X — @askalphaxiv