Google Flow now supports doodles (annotations) for images. Doodles can be used to instruct video models on how the image should be animated.
MULTIMODAL AI
-

Multi-Agent Self-Critique Loop for Automated Video Refinement
By
–
The self-critique loop… After picking the best video, three specialized agents critique it: • Visual agent → checks motion, lighting, focus
• Audio agent → checks sound sync and clarity
• Context agent → checks story flow and coherence Then a “reasoning agent” rewrites -

VISTA: AI-Driven Video Generation and Selection Workflow
By
–
VISTA breaks your idea into scene-by-scene plans with full details camera angles, mood, sounds, transitions. Then it generates multiple versions and picks the best one in a pairwise tournament judged by an MLLM. Think of it like AI video “survival of the fittest.”
-

VISTA improves consistency in generative video models like Veo and Sora
By
–
Text-to-video models like Veo 3 and Sora are incredible but fragile.
Change one word in your prompt and your video falls apart. VISTA fixes that. It doesn’t just generate video, it plans, judges, and iterates like a director reviewing takes on set. -

Google introduces VISTA self-improving video generation agent
By
–
Holy shit…Google just dropped a self-improving video generation agent It’s called VISTA, and it literally rewrites its own prompts to make videos better every single generation. No retraining. No fine-tuning. Just pure test-time self-reflection. Here’s how it works: →
-

LFM2-VL-3B: New Large Vision Language Model Released
By
–
LFM2-VL-3B just dropped! It's a bigger version of our VLMs with fast inference and strong performance. Look at how gracefully it distinguishes dogs from ice cream scoops
-
Grok web gets HD upscale option for Imagine videos
By
–
Grok on the web is getting an option to upscale Imagine videos to HD https://t.co/7UJj5NzRNv pic.twitter.com/bBHD3vo36y
— 🚨 AI News | TestingCatalog (@testingcatalog) 22 octobre 2025Grok on the web is getting an option to upscale Imagine videos to HD
-
Data Mixing Balance: Vision-Language Skills and Regional Imbalance
By
–
That's a great question and will take the opportunity to dash off a bit on the data mixing. Mixing data is a tricky balance it turns out. There were two main factors at play: – we wanted to keep general vision-language skills.
– and had unbalanced regions and languages: think -
Training multilingual cultural model improves multimodal benchmark performance
By
–
Thank you so much and appreciate you taking time to check the paper. We trained a model(
https://
huggingface.co/neulab/Cultura
lPangea-7B
…) on the subset of the dataset and were able to improve performance on many multilingual/cultural multimodal benchmarks. Yeahhh, would be great too if you could also -
Stable Audio 2.5 live chat with ComfyUI demonstrates inpainting
By
–
We joined our friends at @ComfyUI for a live chat about what's new with Stable Audio 2.5 🥁
— Stability AI (@StabilityAI) 21 octobre 2025
Check out CJ Carr from our Audio Research team demoing the model’s inpainting capability.
You can watch the full episode here for more insights and examples 👉 https://t.co/430lf6kve3 pic.twitter.com/2A0KHZieQBWe joined our friends at @ComfyUI for a live chat about what's new with Stable Audio 2.5 🥁 Check out CJ Carr from our Audio Research team demoing the model’s inpainting capability. You can watch the full episode here for more insights and examples 👉 bit.ly/43nRa5x
→ View original post on X — @stabilityai, 2025-10-21 19:31 UTC
