Introducing gpt-realtime — our best speech-to-speech model for developers, and updates to the Realtime API
MULTIMODAL AI
-
Google Nano-Banana AI Model Now Available in Gemini App
By
–
Try nano-banana for yourself in the @GeminiApp now, it's super easy:
-

Real-time Video Generation Model Launches Beta
By
–
today, we're making another step towards the future.
— KREA AI (@krea_ai) 28 août 2025
introducing our first Real-time Video generation model.
join the beta 👇 pic.twitter.com/MyNnRVke9Ptoday, we're making another step towards the future. introducing our first Real-time generation model. join the beta
-
Gemini Flash 2.5 Image Generation and Editing Capabilities Interview
By
–
Great interview with my awesome colleagues @19kaushiks @robertriachi @m__dehghani and @nbrichtova by @OfficialLoganK about the new image generation and editing capabilities in our latest Gemini Flash 2.5 model (also known as nano banana 🍌) https://t.co/vwPAbPpple
— Jeff Dean (@JeffDean) 28 août 2025Great interview with my awesome colleagues @19kaushiks @robertriachi @m__dehghani and @nbrichtova by @OfficialLoganK about the new image generation and editing capabilities in our latest Gemini Flash 2.5 model (also known as nano banana )
-
ChatGPT-5 vs Gemini 2.5 Pro for PDF extraction
By
–
I needed to extract a list of 40 prompt names from a PDF for a client. Used ChatGPT-5 Thinking, wasted an hour. Used Gemini 2.5 Pro, it did it in 1 minute. Choose your tool wisely.
-
Creative video generation workflow using AI prompting
By
–
Nano banana pour la première Frame, ensuite JSON prompting avec ChatPPT et veo 3 pour la vidéo
-

Microsoft VibeVoice-1.5B: Open-Source Multilingual Speech Generation
By
–
🎙️ Microsoft just dropped VibeVoice-1.5B — an open-source speech model that can:
— Futurepedia – Learn to Leverage AI (@futurepedia_io) 28 août 2025
▫️Generate 90 minutes of audio
▫️Up to 4 unique voices
▫️Expressive
▫️Multilingual
▫️MIT-licensed.
A game-changer for voice AI.
👉 For more AI breakthroughs, follow @futurepedia_io
Swing by… pic.twitter.com/5lZ1weu6feMicrosoft just dropped VibeVoice-1.5B — an open-source speech model that can: Generate 90 minutes of audio Up to 4 unique voices
Expressive
Multilingual
MIT-licensed. A game-changer for voice AI. For more AI breakthroughs, follow @futurepedia_io Swing by -

Google DeepMind Perch v2 Identifies Threatened Species from Audio
By
–
And it continues at Google! Definitely… Google DeepMind has just launched Perch v2, an open-source AI model capable of analyzing nature sounds to identify threatened species from audio recordings. Here's how it's already been used: – Discovery of a new population of the rare
-

Tencent Launches HunyuanVideo-Foley Open-Source Audio Generation AI
By
–
Tencent launched HunyuanVideo-Foley, an open-source Text-Video-to-Audio AI
— The Rundown AI (@TheRundownAI) 28 août 2025
Trained on 100k-hours of data, it generates contextually-aware soundscapes for scenes ranging from natural landscapes to animated shorts
Also achieves SOTA on multiple benchmarkspic.twitter.com/0CWII2KWyeTencent launched HunyuanVideo-Foley, an open-source Text-Video-to-Audio AI Trained on 100k-hours of data, it generates contextually-aware soundscapes for scenes ranging from natural landscapes to animated shorts Also achieves SOTA on multiple benchmarks
-

Google Upgrades Vids Platform with AI Avatars and Video Generation
By
–
Google upgraded its 'Vids' AI video editing platform with new features, including:
— The Rundown AI (@TheRundownAI) 28 août 2025
—AI avatars to deliver content as videos (like Synthesia)
—Veo-3 powered image-to-video to animate product shots and more
—Automatic transcript cleanup to cut filler wordspic.twitter.com/xy3fnzx6eEGoogle upgraded its 'Vids' AI video editing platform with new features, including: —AI avatars to deliver content as videos (like Synthesia)
—Veo-3 powered image-to-video to animate product shots and more
—Automatic transcript cleanup to cut filler words