LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding LLaVA-ST is a multimodal large language model designed for spatial-temporal understanding in videos, addressing challenges in coordinate alignment and feature compression. Problem:
LLaVA-ST: Multimodal LLM for Spatial-Temporal Video Understanding
By
–
