Learn more about our work with @gopuff to build a personalized shopping assistant with chat, voice, and image models https://
x.ai/news/grok-gopu
ff
…
MULTIMODAL AI
-

xAI and Gopuff build personalized shopping assistant with chat, voice, image models
By
–
-
Agent verifies brand logos and details in ads for $5–$8
By
–
The agent > locks verified brand details, logos, and claims as constraints > then checks every frame against them with its own vision intelligence to catch logo drift and product errors before delivery. A finished ad runs about $5 to $8 and returns dozens of
-
Language Swap by Pika MCP lets you speak any language in videos
By
–
Take your content global, already!
— Pika (@pika_labs) 9 juin 2026
The Language Swap skill via Pika MCP swaps the language you’re speaking in any video, making it look and sound like you speak fluent…anything. pic.twitter.com/kApd9yEbvXTake your content global, already! The Language Swap skill via Pika MCP swaps the language you’re speaking in any video, making it look and sound like you speak fluent…anything.
-
HyperFrames engine becomes Claude connector for easy AI video
By
–
The HyperFrames engine leaving the terminal and becoming a Claude connector is a bigger deal than it looks.
— Chubby♨️ (@kimmonismus) 9 juin 2026
Ask for a video the way you'd ask for the report. No repo, no setup. That's the version of AI video that non-developers will actually use. https://t.co/Z8q8tb0a7pThe HyperFrames engine leaving the terminal and becoming a Claude connector is a bigger deal than it looks. Ask for a video the way you'd ask for the report. No repo, no setup. That's the version of AI video that non-developers will actually use.
-

Claude 5 Fable: state-of-the-art on benchmarks, excels on longer and complex tasks
By
–



Claude 5 Fable tl;dr – It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research -The longer and more complex the task, the larger Fable 5’s lead over our other
-
Gemini 3.5 Live Translate: Google’s new voice translation model
By
–
Speech translation has been one of the longest-running ML efforts at Google, and we’ve come a long way. Gemini 3.5 Live Translate is our latest speech-to-speech model, supporting 70+ languages. It enables more natural conversations across languages in everyday products and… pic.twitter.com/ZqzSW3V8lY
— Jeff Dean (@JeffDean) 9 juin 2026Voice translation has been one of the most enduring efforts in machine learning at Google, and we have come a long way. Gemini 3.5 Live Translate is our latest speech-to-speech translation model, supporting over 70 languages. It enables
-
Gemini 3.5 Live Translate: spoken translation, multi-speaker, no Klingon
By
–
Gemini 3.5 Live translate: Stream in speech, and stream out the spoken translation.
— fofr (@fofrAI) 9 juin 2026
It also magically works with multiple speakers.
It does not work with Klingon (I tried).
Try it on AI Studio:https://t.co/pTVEPZ0llf
pic.twitter.com/FMtlj5G2gbGemini 3.5 Live Translate: Speech stream, and spoken translation stream. It also works magically with multiple speakers. It does not work with Klingon (I tried). Try it on AI Studio: https:// aistudio.google.com/live?model=gem ini-3.5-live-translate-preview …
-
Latent Spatial Memory for Video World Models
By
–
Latent Spatial Memory for Video World Models pic.twitter.com/sJIpofmrmQ
— AK (@_akhaliq) 9 juin 2026Latent Spatial Memory for World Models
-
Evaluation of interaction and spatial reasoning of multimodal agents
By
–
SpatialWorld Evaluation of interaction and spatial reasoning of multimodal agents in real-world tasks
-
Introducing Gemini 3.5 Flash Live Translate with 70+ languages
By
–
Introducing Gemini 3.5 Flash Live Translate, our real time speech to speech translation model which supports more than 70 languages (both in and out), and is so natural.
— Logan Kilpatrick (@OfficialLoganK) 9 juin 2026
It is available in the Gemini API, AI Studio, & Google Translate right now + coming soon to Google Meet!! pic.twitter.com/nDiaqrbHgFIntroducing Gemini 3.5 Flash Live Translate, our real time speech to speech translation model which supports more than 70 languages (both in and out), and is so natural. It is available in the Gemini API, AI Studio, & Google Translate right now + coming soon to Google Meet!!