DeepPHY Benchmarking Agentic VLMs on Physical Reasoning
MULTIMODAL AI
-
Runway Aleph Enables Granular Object Control in Video Generation
By
–
Runway Aleph has granular object control. Which means you can add to or alter your video in ways which feel both natural and realistic without any complex prompting or key framing. Or, you can break the laws of physics altogether. All you need to do is tell Aleph what you want. pic.twitter.com/HMfi4rArjS
— Runway (@runwayml) 8 août 2025Runway Aleph has granular object control. Which means you can add to or alter your video in ways which feel both natural and realistic without any complex prompting or key framing. Or, you can break the laws of physics altogether. All you need to do is tell Aleph what you want.
-

AI Video Editing: Consistent Modifications Without Unintended Side Effects
By
–
Using AI, you can modify parts of videos without the annoying side effect of subtly changing everything else. Lighting effects (as seen on the shirt) are very consistent. pic.twitter.com/RsROwR3agJ
— Varun Mayya (@waitin4agi_) 8 août 2025Using AI, you can modify parts of videos without the annoying side effect of subtly changing everything else. Lighting effects (as seen on the shirt) are very consistent.
-
AI cameras now decide what happens at weddings
By
–
There was a time the camera captured what happened at the wedding. Now the camera decides what happens at the wedding.
-

ChatGPT-5 Released: PhD-Level Reasoning and Multimodal Capabilities
By
–
🚨BIG AI NEWS ALERT!
— Futurepedia – Learn to Leverage AI (@futurepedia_io) 8 août 2025
It’s official – ChatGPT-5 has just dropped… and it’s a TOTAL game-changer!
Here’s what’s new:
🔹PhD-level reasoning – handles complex problems with far greater accuracy.
🔹Next-gen multimodal skills – works seamlessly with text, images, voice, and video.… pic.twitter.com/dWsyLoK3CJBIG AI NEWS ALERT! It’s official – ChatGPT-5 has just dropped… and it’s a TOTAL game-changer! Here’s what’s new:
PhD-level reasoning – handles complex problems with far greater accuracy.
Next-gen multimodal skills – works seamlessly with text, images, voice, and video. -
GPT-5 Creates Audio-Reactive Breathing Mandala Single Shot
By
–
3. GPT-5 single shot creates Breathing Mandala with audio reactivity.pic.twitter.com/mALeP4kWms
— Shubham Saboo (@Saboo_Shubham_) 8 août 20253. GPT-5 single shot creates Breathing Mandala with audio reactivity.
-
GPT-5 Launch: 10 Insane Use Cases Already Built
By
–
GPT-5 just dropped less than 9 hours ago.
— Shubham Saboo (@Saboo_Shubham_) 8 août 2025
And people are already building wild stuff with it.
Here are 10 insane use cases I’ve seen so far:
1. Snake game in a Hexagon pic.twitter.com/qfws7RvFX2GPT-5 just dropped less than 9 hours ago. And people are already building wild stuff with it. Here are 10 insane use cases I’ve seen so far: 1. Snake game in a Hexagon
-
GPT-5 Review Series: Multi-part Analysis and Taste Test
By
–
part 1: https://
latent.space/p/gpt-5-review
part 2: https://
latent.space/p/gpt5-router
part 3: https://
latent.space/p/gpt5-vision we will be publishing the taste test in part 4 -

Router Layer Latency Impact on Vision Processing Tasks
By
–
another consequence of knowing there's a router layer is being able to understand the latency delta between the raw thinkies. in the case of hard vision inputs, it adds 2-3s latency on average.
-

Gemini adds multimedia to responses
By
–
ICYMI: Gemini changelog got updated with new entries and has been moved to a new home! "Gemini now provides a richer learning experience by automatically integrating high-quality images, diagrams, and YouTube videos directly into its responses"
