I just read this paper called "Chain-of-Visual-Thought (COVT)" and it basically teaches VLMs to see and think at the same time not in text, but in continuous visual tokens. Here’s the wild part: Instead of forcing models to reason through words (which destroys all the
