Why do AI image generators keep forgetting your detailed descriptions? A team from Fudan University, Alibaba, and Baidu introduces Prompt Reinjection. They found that multimodal diffusion transformers gradually lose prompt info in deeper layers. Their training-free fix:
MULTIMODAL AI
-
Pipeline for VQA data from Wikidata with LLM/VLM filters
By
–
A pipeline for creating VQA data using knowledge bases (wikidata). We demonstrated the method on world/cultural knowledge and could be extended to other domains. (there are rephrasing/filtration steps that require LLMs/VLMs, but the released data does not).
-

Data2Story: AI framework transforms raw data into verifiable news articles
By
–
/3 Data2Story: an AI framework transforming raw data into verifiable, multimodal news articles. Data2Story orchestrates a virtual newsroom of specialized AI agents to analyze data, design multimedia assets, and write narratives. Crucially, it features an "Inspector" that
-
AI provides practical use cases for athletes, creators, and low vision
By
–
The use cases are not just novelty. They are practical: → Athletes getting live performance feedback
→ Creators capturing POV content instantly
→ People with low vision receiving real time descriptions of their environment When AI can understand your surroundings, it stops -
Glasses as AI interface: see, understand, respond in real time
By
–
AI’s next interface won’t be another app.
— Ronald van Loon (@Ronald_vanLoon) 22 juin 2026
It will be your eyes.
That sounds futuristic, but it is already becoming real with glasses that can see what you see, understand context, and respond in real time.
Here’s why that matters… pic.twitter.com/GFm7VRhHDVAI’s next interface won’t be another app. It will be your eyes. That sounds futuristic, but it is already becoming real with glasses that can see what you see, understand context, and respond in real time. Here’s why that matters…
-
xAI launches Grok Imagine Video 1.5 with integrated audio
By
–
4/ xAI shipped Grok Imagine Video 1.5 (GA).
— Futurepedia – Learn to Leverage AI (@futurepedia_io) 22 juin 2026
Image-to-video with sound effects, ambience and dialogue generated in the same pass — no separate audio step.
The Fast variant: ~25 sec for a 6-second 720p clip, down from 40+.
Live on the API, https://t.co/dvlgFY9jys, iOS, Android. pic.twitter.com/kwOQ3mZFiM4/ xAI shipped Grok Imagine 1.5 (GA). Image-to-video with sound effects, ambience and dialogue generated in the same pass — no separate audio step. The Fast variant: ~25 sec for a 6-second 720p clip, down from 40+. Live on the API, http://
grok.com, iOS, Android. -

Google launches AI video and music creation on Android phones
By
–
Your Android phone can now generate videos and music with AI. Google just launched: Gemini Omni → create and edit videos through conversation Lyria 3 → generate original songs from a prompt or photo The biggest shift? Creative tools that once required specialized
-

Do as I Do: Dexerous Manipulation Data from Everyday Human Videos
By
–
“Do as I Do: Dexterous Manipulation Data from Everyday Human Videos” With how robot dexterity is bottlenecked by data as teleoperation and MoCap are expensive and internet videos are only observational, this paper turns normal RGB human videos into executable robot hand
-
Claude AI self-optimizes upon discovering voice harness
By
–
I met a guy last night building a next-level voice harness. He hooked Claude up to it and it realized quickly the way the voice model worked and optimized itself for that use.
