Voice agents are so back!! Today we’re launching 3 new realtime audio models in the API: GPT-Realtime-2
GPT-5-class reasoning for voice agents that can use tools, recover from interruptions, and carry longer conversations with 128K context GPT-Realtime-Translate
Live
MULTIMODAL AI
-

New realtime audio models launched
By
–
-
ElevenLabs introduces Studio Agent AI co-editor in ElevenCreative
By
–
Introducing Studio Agent in ElevenCreative.
— ElevenLabs (@ElevenLabs) 7 mai 2026
Studio is the timeline editor where creators and marketers mix voiceovers, music, sound effects, and video into finished content. Now it has an AI co-editor built in. pic.twitter.com/FXeO3k6abcIntroducing Studio Agent in ElevenCreative. Studio is the timeline editor where creators and marketers mix voiceovers, music, sound effects, and video into finished content. Now it has an AI co-editor built in.
-
AI agent places sound effects frame-by-frame in your video
By
–
See it lay out your video exactly as you imagined. The agent analyzes your clips frame by frame – place a sound effect when your product is revealed, or add a swoosh at the exact moment text appears.
-
AI Co-Editor Brings Your Video Idea to Life
By
–
Now, Studio Agent adds an AI co-editor to this workflow. Describe your idea, and watch the agent bring it to life. It asks what you need – video length, tone, structure – then builds a first draft on the timeline.
-
ElevenLabs Studio: AI voiceovers, music, and audio editing
By
–
Our Studio product lets you:
– Add voiceovers from 10,000+ voices in 32+ languages.
– Generate background music and sound effects on the timeline.
– Edit spoken audio by editing the script – no re-recording.
– Upload video and enhance it with AI audio. -

Discussion on voice cloning AI capabilities and real-time models
By
–
Donnez-moi encore un peu de temps, je vais vous faire une démo avec certains des derniers modèles qui promettent de faire du clonage en quelques secondes. Entre TTS, speech-to-speech, temps réel ou non, il y a un monde. Vous allez voir qu’on est loin de ce que ce reportage
-
Casual AI audio generation and compression workflow
By
–
Essaye et met le résultat avec 3sc d'audio compressé en francais 😉 .
-
AI video character consistency breakthrough
By
–
1/
— Charly Wargnier (@DataChaz) 7 mai 2026
Let’s tackle the biggest headache in AI video right now:
Character consistency.
Watch this sequence.
A character moves through a subway, a convenience store, a diner, and an overpass.
Her face, proportions, and identity stay 100% locked the entire time.
That’s the… pic.twitter.com/y9tguam8601/ Let’s tackle the biggest headache in AI video right now: Character consistency. Watch this sequence. A character moves through a subway, a convenience store, a diner, and an overpass. Her face, proportions, and identity stay 100% locked the entire time. That’s the
-
Fundamental interface shift: from ask AI to AI understands everything around you
By
–
We’re moving from “ask AI anything” → to “AI understands everything around you.” That’s a fundamental interface shift. If you’re building products, this changes UX, data strategy, and distribution. I break this down in this latest video with Meta. Curious how you see this
-
Multimodal reasoning: AI sees, understands, decides instantly
By
–
The real leap isn’t better answers. It’s multimodal reasoning. → AI that sees images
→ Understands context
→ Makes decisions instantly Example: Point your phone at food → it suggests the healthiest option
Scan a product → it gives a real comparison, not biased reviews