AI Dynamics

Global AI News Aggregator

About

Chain-of-Modality: New Prompting Strategy for Vision-Language Models

DeepMind and Stanford just dropped a new paper!
They propose Chain-of-Modality, a prompting strategy that enables Vision-Language Models to reason over multimodal human demonstrations — bridging modalities step by step.
A smart way to unlock richer, more grounded reasoning.

→ View original post on X — @jiqizhixin