Small update today. Many requested a "anti-prompting" feature for V8 models (which existed in previous versions) which we call the –no flag. This is now available today in V8.1! So if you're trying to get something out of your images (like people) try –no people. Have fun!
MULTIMODAL AI
-
Building ‘Gemini for Science’ with the scientific community
By
–
We're building Gemini for Science with and for the scientific community. In collaboration with 100+ institutions and a trusted tester community that ranges from PhD students to Nobel laureates, we want to make sure this tech is responsible and rigorous enough to tackle real-world
-
New Course on Building Self-Iterating AI Agents for Media Generation
By
–
New course: Build AI agents that generate images and videos — an under-explored frontier. A key to performance is having the agent evaluate its own output, and iterate to improve quality. This short course is built together with @googlecloudtech and taught by Katie Nguyen and… pic.twitter.com/xxuEvZs8HP
— Andrew Ng (@AndrewYNg) 20 mai 2026New course: Build AI agents that generate images and videos — an under-explored frontier. A key to performance is having the agent evaluate its own output, and iterate to improve quality. This short course is built together with @googlecloudtech and taught by Katie Nguyen and
-

ESI-Bench: Towards Embodied Spatial Intelligence and Perception-Action Loops
By
–
ESI-Bench Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
-
Gemini Shifts from Chatbot to Autonomous Task-Oriented AI System
By
–
6/ Gemini is no longer a chatbot. It is now a system that runs tasks on its own. Spark in the background. Antigravity for development. Omni Flash for media. The chat box is the smallest piece of the product.
-
Breakdown of Gemini 3.5 Flash architecture and its operational layers
By
–
5/ Gemini 3.5 Flash is the core model. The other three sit on top of it.
Omni Flash handles video. Spark runs your tasks in the background. Antigravity is the environment developers use to build with it. One model, three different surfaces. -
Gemini Omni Model Released for Multimodal Video Creation and Editing
By
–
2/ Gemini Omni.
— AI Highlight (@AIHighlight) 20 mai 2026
A new model that creates and edits video from any input. Text, audio, image, or another video.
Live in the Gemini app, Flow, and YouTube Shorts.
Use for: video creation and editing. pic.twitter.com/AO0mldizDr2/ Gemini Omni. A new model that creates and edits video from any input. Text, audio, image, or another video. Live in the Gemini app, Flow, and YouTube Shorts. Use for: video creation and editing.
-
Google Launches Gemini 3.5 Flash and New AI Products
By
–
🚨Breaking: Google just launched 4 new Gemini products at I/O 2026 yesterday.
— AI Highlight (@AIHighlight) 20 mai 2026
Gemini 3.5 Flash. Gemini Omni. Gemini Spark. Antigravity 2.0.
Here is what each one is, what it does, and when to use it: pic.twitter.com/xjV2ya1BvVBreaking: Google just launched 4 new Gemini products at I/O 2026 yesterday. Gemini 3.5 Flash. Gemini Omni. Gemini Spark. Antigravity 2.0. Here is what each one is, what it does, and when to use it:
-
Stability AI releases Stable Audio 3.0
By
–
Start experimenting with Stable Audio 3.0: Small + Medium → @huggingface https://
huggingface.co/collections/st
abilityai/stable-audio-3
… Large → Stability AI API or self-host with an enterprise license https://
stability.ai/stable-audio -

Comparison of the Stable Audio 3.0 Model Family
By
–
Compare the Stable Audio 3.0 family, four new models designed for different use cases and deployment options: