"Benchmarking Visual State Tracking in Multimodal Understanding" A new benchmark for tracking visual states. Even though video MLLMs can describe clips really well, they still cannot reliably track what changes over time. This benchmark contains 834 videos and 1,500
New Benchmark for Visual State Tracking in Video Understanding
By
–
