How well can AI read the room? R³-VQA is a new video QA benchmark pushing LVLMs to reason like humans in complex social scenes—tracking beliefs, emotions, intentions, and more. Turns out, even top models still struggle with consistent Theory of Mind.
R³-VQA Benchmark Tests AI Theory of Mind in Social Scenes
By
–