Top stories in AI today: – Sam Altman on Dev Day, AGI, and more
– Google releases Gemini 2.5 Computer Use
– Create LinkedIn carousels in ChatGPT with Canva
– Duke’s AI for smarter drug delivery
– 4 new AI tools, community workflows, and more Read more: https://
therundown.ai/p/exclusive-in
terview-sam-altman-on-dev-day-and-ais-future
…
MULTIMODAL AI
-

Sam Altman, Gemini 2.5, and AI tools roundup
By
–
-

Panoramic Vision Survey: Bridging Gap Between Perspective Views
By
–
#PapersAccepted by Jiqizhixin
Our report: https://
mp.weixin.qq.com/s/O-4L9pACS-kG
xTX6fCLtOA
… One Flight Over the Gap: A Survey from Perspective to Panoramic Vision Insta360 Research, University of California, San Diego, and others
Project: https://
insta360-research-team.github.io/Survey-of-Pano
rama/
…
Paper: https://
arxiv.org/pdf/2509.04444
Repo: -
Teaching AI Panoramic Vision: Survey of 300+ Works
By
–
How do we teach AI to see the world in 360°?
— 机器之心 JIQIZHIXIN (@jiqizhixin) 8 octobre 2025
A new survey dives deep into panoramic vision — exploring how models adapt from standard perspective images to omnidirectional images (ODIs) used in VR, robotics, and autonomous driving.
The paper reviews 300+ works and identifies… pic.twitter.com/tgmWmnRwJIHow do we teach AI to see the world in 360°? A new survey dives deep into panoramic vision — exploring how models adapt from standard perspective images to omnidirectional images (ODIs) used in VR, robotics, and autonomous driving. The paper reviews 300+ works and identifies
-

voice-chat-03 Multimodal Chat UI with ElevenLabs Agent
By
–
☑️ voice-chat-03
— Futurepedia – Learn to Leverage AI (@futurepedia_io) 8 octobre 2025
💬 Next-gen multimodal chat is here.
voice-chat-03 packs audio, text, and visual interactions into one sleek UI — complete with built-in state management.
Just pass your ElevenLabs Agent ID as a prop, and you’re ready to ship. pic.twitter.com/78vpUCUeWKvoice-chat-03 Next-gen multimodal chat is here. voice-chat-03 packs audio, text, and visual interactions into one sleek UI — complete with built-in state management. Just pass your ElevenLabs Agent ID as a prop, and you’re ready to ship.
-

Google Launches Multimodal Computer Use Model Outperforming Sonnet
By
–
Nuevo modelo de Google de uso de ordenador (multimodalidad de visión y acciones en GUI con clicks, teclado, etc). Se coloca por delante de Sonnet 4.5 y OpenAI Agent y además con una latencia más baja, que para agilizar la navegación se agradece!
-

GPT-1 Image vs Mini: Price and Performance Comparison
By
–
Which one is which: gpt-1-image or the new gpt-1-image-mini? Can you guess? The mini version is 5x cheaper than the full model: $0.05 vs $0.25 per image. 5 examples below. Models A and B are consistent across all examples. Answer in the last post. 1) "Ultra-macro photograph of
-
Grok Imagine needs significant improvements in output quality
By
–
Grok Imagine output is pretty bad. It needs a lot of improvement.
-
Optimizing 2x Speed Video Generation with Frame Rate Tradeoffs
By
–
Thanks Matthew! I really want to get this working super well. My next step is trying to get it to generate 2x speed videos… frame rate takes a hit but it may enable longer videos w/ fewer jarring transitions
-

Grok Imagine AI: Share Your Feedback for Model Improvements
By
–
We’d love to hear your feedback on Grok Imagine. The team will be incorporating your responses into continuous model upgrades, so please be vocal on X.
— xAI (@xai) 7 octobre 2025
Try it here: https://t.co/2DPEzEZ03e pic.twitter.com/fihPWOASrfWe’d love to hear your feedback on Grok Imagine. The team will be incorporating your responses into continuous model upgrades, so please be vocal on X. Try it here: https://
grok.com/imagine -

Expressive singing brought to life by AI with synchronized emotion
By
–
It also brings expressive singing to life with clear vocals and synchronized emotion. pic.twitter.com/xY5DoBmtVI
— xAI (@xai) 7 octobre 2025It also brings expressive singing to life with clear vocals and synchronized emotion.