API docs here: https://
ai.google.dev/gemini-api/doc
s/image-generation#video-to-image
… For a YouTube video, if your video is long, scale down the `fps` to keep the number of frames within the context window.
MULTIMODAL AI
-

Gemini API Video-to-Image Generation: Frame Optimization
By
–
-
Nano Banana 2 Video-to-Image Feature Launch
By
–
Fun new feature landing in Nano Banana 2 today:
— fofr (@fofrAI) 28 mai 2026
video-to-image
Pass a video or YouTube URL to Nano Banana 2 along with your prompt, and it'll use the video as context to make the image.
> create a comic strip from this video pic.twitter.com/1raoUB1UlTFun new feature landing in Nano Banana 2 today: video-to-image Pass a video or YouTube URL to Nano Banana 2 along with your prompt, and it'll use the video as context to make the image. > create a comic strip from this video
-
Generative Supervision for Embodied Intelligence
By
–
GEM
— AK (@_akhaliq) 28 mai 2026
Generative Supervision Helps Embodied Intelligence pic.twitter.com/IlGPbxkwHSGEM Generative Supervision Helps Embodied Intelligence
-
Dubbing v2 Preserves Tone via Direct Performance Conditioning
By
–
Dubbing v2 fixes flat audio by conditioning directly on the original performance, not a transcript. It ensures that tone, emotion, and delivery are preserved. This is the problem AI dubbing had never previously solved.
-
AI Dubbing v2: Multilingual Sync-Aware Translation
By
–
Dubbing v2 also adapts phrasing for natural delivery across 90+ languages. Sync-aware translation logic means that starts and stops align with the original.
-

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
By
–
Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
-

Masked Region Transformer for Layered Image Generation
By
–
MRT Masked Region Transformer for Layered Image Generation and Editing at Scale
-
104M Image-Text Pair Dataset Released on Hugging Face
By
–
With 104M of image-text pairs, this is one of the largest, if not the largest, openly-licensed image dataset
— Julien Chaumond (@julien_c) 28 mai 2026
And it's on @huggingface!!
Kudos @heyjasperai https://t.co/mTwGfZUzZUWith 104M of image-text pairs, this is one of the largest, if not the largest, openly-licensed image dataset And it's on @huggingface
!! Kudos @heyjasperai -
Creator builds drawing-capture tool with Google Flow and Gemini Omni
By
–
WOW
— Charly Wargnier (@DataChaz) 28 mai 2026
this guy literally vibe-coded his own drawing-capture tool using Google Flow, then asked Gemini Omni for photorealistic red yarn, and created absolute MAGIC 🤯 pic.twitter.com/7VV2bKnzkfWOW this guy literally vibe-coded his own drawing-capture tool using Google Flow, then asked Gemini Omni for photorealistic red yarn, and created absolute MAGIC
-

Meta Ray-Ban Glasses Footage Used to Train AI
By
–
BREAKING: Meta's Ray-Ban smart glasses record what you see. Some of that footage is watched by people. When a user shares data to improve Meta AI, the clips their glasses captured can be sent to human contractors who review and label them by hand. The reviewers work for a firm