The FAQ doesn't help clarify which underlying model is being used: https://
help.openai.com/en/articles/84
00625-voice-mode-faq
…
GENERATIVE AI
-

OpenAI Voice Mode FAQ Lacks Clarity on Underlying Model
By
–
-
iPhone ChatGPT Voice Mode: GPT-4o versus gpt-realtime models
By
–
Does anyone understand the relationship between iPhone ChatGPT voice mode and the underlying models? Is it still stuck with the old GPT-4o voice model, or is it using the new gpt-realtime model? How does gpt-realtime relate to GPT-5?
-

GLM-4.5 Surpasses Claude Opus in Tool Use Efficiency
By
–
Z AI's GLM-4.5 just dethroned Claude Opus 4 in tool use with nearly 100x lower costs The model scored 70.85 on the Berkeley Function-Calling Leaderboard at just $2.9, while Opus 4 scored 70.36 at a cost of $207.12
-
Tencent Launches Hunyuan World Model and Multilingual Translation System
By
–
Tencent launched two new Hunyuan models:
— The Rundown AI (@TheRundownAI) 2 septembre 2025
—HunyuanWorld-Voyager, an ultra-long-range world model with native 3D reconstruction and memory
—Hunyuan-MT-7B and Hunyuan-MT-Chimera-7B, a joint AI translation system that outperforms rivals across 33 languagespic.twitter.com/NAeOdjGQa4Tencent launched two new Hunyuan models: —HunyuanWorld-Voyager, an ultra-long-range world model with native 3D reconstruction and memory
—Hunyuan-MT-7B and Hunyuan-MT-Chimera-7B, a joint AI translation system that outperforms rivals across 33 languages -

Microsoft Launches 10B VibeVoice Text-to-Speech Model
By
–
Microsoft just launched a larger 10B parameter version of its VibeVoice text-to-speech model
— The Rundown AI (@TheRundownAI) 2 septembre 2025
Available under MIT license, it generates multi-speaker podcasts (going up to 45 minutes) in minutes
Supports up to 32K contextpic.twitter.com/fO9JOT5RkJMicrosoft just launched a larger 10B parameter version of its VibeVoice text-to-speech model Available under MIT license, it generates multi-speaker podcasts (going up to 45 minutes) in minutes Supports up to 32K context
-

Apple Launches FastVLM and MobileCLIP2 Open Models
By
–
Apple launched two new open models: FastVLM and MobileCLIP2
— The Rundown AI (@TheRundownAI) 2 septembre 2025
The models are up to 85x faster and 3.4x smaller than previous work
Suitable for high-res image processing tasks like OCR, image captioning, visual question answering, and image understandingpic.twitter.com/uG9ObsWACuApple launched two new open models: FastVLM and MobileCLIP2 The models are up to 85x faster and 3.4x smaller than previous work Suitable for high-res image processing tasks like OCR, image captioning, visual question answering, and image understanding
-
Apple Launches Open-Source Real-Time Vision Language Models
By
–
AI NEWS: Apple just launched two open-source models for real-time vision language applications. Plus, more news from Microsoft, Tencent, Z AI, and OpenAI. Here's everything you need to know:
-
Moving Beyond Basic ChatGPT to Agent Building
By
–
I've been stuck using ChatGPT for basic tasks. Learning to build actual agents could be a game changer.
-

Google DeepMind August Releases: Gemini, Veo, Imagen and More
By
–
The craziest Google I/O to date! In August at Google DeepMind, it was: Nano Banana (Gemini 2.5 Flash Image)
Gemini Embedding
Veo 3 Fast
Genie 3
Imagen 4 Fast
Gemma 3 270M
Perch 2
Kaggle Game Arena
Gemini API Url Context
AI Studio Builder (UI overhaul, prompt suggestions, GitHub -
GPT-Realtime Model Supports Video Input via Image Frames
By
–
The gpt-realtime model supports image input, my guess is that you can feed it a video frame once per second or more as a static image