For August it's mainly the release and impact of GPT-5, plus the new wave of image editing multi-modal LLMs represented by Qwen-Image-Edit (open weights) and Gemini Nano Banana Here's the August newsletter table of contents
MULTIMODAL AI
-
Tencent Launches Hunyuan World Model and Multilingual Translation System
By
–
Tencent launched two new Hunyuan models:
— The Rundown AI (@TheRundownAI) 2 septembre 2025
—HunyuanWorld-Voyager, an ultra-long-range world model with native 3D reconstruction and memory
—Hunyuan-MT-7B and Hunyuan-MT-Chimera-7B, a joint AI translation system that outperforms rivals across 33 languagespic.twitter.com/NAeOdjGQa4Tencent launched two new Hunyuan models: —HunyuanWorld-Voyager, an ultra-long-range world model with native 3D reconstruction and memory
—Hunyuan-MT-7B and Hunyuan-MT-Chimera-7B, a joint AI translation system that outperforms rivals across 33 languages -

Microsoft Launches 10B VibeVoice Text-to-Speech Model
By
–
Microsoft just launched a larger 10B parameter version of its VibeVoice text-to-speech model
— The Rundown AI (@TheRundownAI) 2 septembre 2025
Available under MIT license, it generates multi-speaker podcasts (going up to 45 minutes) in minutes
Supports up to 32K contextpic.twitter.com/fO9JOT5RkJMicrosoft just launched a larger 10B parameter version of its VibeVoice text-to-speech model Available under MIT license, it generates multi-speaker podcasts (going up to 45 minutes) in minutes Supports up to 32K context
-

Apple Launches FastVLM and MobileCLIP2 Open Models
By
–
Apple launched two new open models: FastVLM and MobileCLIP2
— The Rundown AI (@TheRundownAI) 2 septembre 2025
The models are up to 85x faster and 3.4x smaller than previous work
Suitable for high-res image processing tasks like OCR, image captioning, visual question answering, and image understandingpic.twitter.com/uG9ObsWACuApple launched two new open models: FastVLM and MobileCLIP2 The models are up to 85x faster and 3.4x smaller than previous work Suitable for high-res image processing tasks like OCR, image captioning, visual question answering, and image understanding
-
Apple Launches Open-Source Real-Time Vision Language Models
By
–
AI NEWS: Apple just launched two open-source models for real-time vision language applications. Plus, more news from Microsoft, Tencent, Z AI, and OpenAI. Here's everything you need to know:
-

Google DeepMind August Releases: Gemini, Veo, Imagen and More
By
–
The craziest Google I/O to date! In August at Google DeepMind, it was: Nano Banana (Gemini 2.5 Flash Image)
Gemini Embedding
Veo 3 Fast
Genie 3
Imagen 4 Fast
Gemma 3 270M
Perch 2
Kaggle Game Arena
Gemini API Url Context
AI Studio Builder (UI overhaul, prompt suggestions, GitHub -
GPT-Realtime Model Supports Video Input via Image Frames
By
–
The gpt-realtime model supports image input, my guess is that you can feed it a video frame once per second or more as a static image
-
Prediction: AI with text-to-speech and lip-sync on pre-recorded
By
–
I predict: Their AI is text-to-speech on images of AI-generated characters, animated in video, with lip-sync pasted from pre-recorded responses, and they will pretend it's live.
-
Apple Releases FastVLM and MobileCLIP2 Vision Models
By
–
If you think @Apple is not doing much in AI, you're getting blindsided by the chatbot hype and not paying enough attention!
— clem 🤗 (@ClementDelangue) 1 septembre 2025
They just released FastVLM and MobileCLIP2 on @huggingface. The models are up to 85x faster and 3.4x smaller than previous work, enabling real-time vision… pic.twitter.com/jYCPukNuiKIf you think @Apple is not doing much in AI, you're getting blindsided by the chatbot hype and not paying enough attention! They just released FastVLM and MobileCLIP2 on @huggingface
. The models are up to 85x faster and 3.4x smaller than previous work, enabling real-time vision -
Create stunningly cinematic voices from text with MiniMax Audio
By
–
Create stunningly cinematic voices
— God of Prompt (@godofprompt) 1 septembre 2025
All from a simple text prompt.
Try MiniMax Audio yourself.
Start with 10,000 free credits. pic.twitter.com/kJsZ2RX9cPCreate stunningly cinematic voices
All from a simple text prompt. Try MiniMax Audio yourself. Start with 10,000 free credits.
