What if your text prompts could generate 3D scenes that actually obey gravity and don’t clip through each other? Carnegie Mellon, HKU, HKUST, and Genesis AI present PAT3D. It combines vision-language models with a physics simulator to arrange objects into stable,
MULTIMODAL AI
-

Doc-V* Agent Introduces OCR-Free Efficient Document Retrieval
By
–
Struggling with AI that can't handle long, multi-page documents? Researchers from HUST and Xiaomi introduce Doc-V*, an OCR-free agent that starts with a thumbnail overview, then actively retrieves and reads only the relevant pages, storing evidence in a working memory. It
-

SecretaryMEITY: Orchestration Layer for AI in Indian Healthcare
By
–
In a panel discussion on Building AI for Indian Healthcare at the AB PM-JAY Auto-Adjudication Hackathon Showcase 2026, @SecretaryMEITY spoke about the need for an orchestration layer that leverages multiple model architectures — LLMs, SLMs and VLMs — to address critical
-
Abacus AI Studio launches agentic video and image creation
By
–
🚨 Abacus AI Studio Offers Agentic Video And Image Capabilities
— Abacus.AI (@abacusai) 8 mai 2026
Use top image and video models including
– Nano Banana Pro
– Sea Dance 2.0
– Kling Motion Control
Using agentic loops powered by Opus 4.7 and GPT 5.5 to create marketing videos and images pic.twitter.com/lJpSw9eeMVAbacus AI Studio Offers Agentic And Image Capabilities Use top image and video models including – Nano Banana Pro
– Sea Dance 2.0
– Kling Motion Control Using agentic loops powered by Opus 4.7 and GPT 5.5 to create marketing videos and images -

Automated Data Engine Trains AI for 3D Scene Understanding from Video
By
–
What if unlabeled YouTube videos could teach AI to understand 3D scenes? BIGAI & collaborators built an automated data engine that extracts 3D training data from raw internet videos — no manual labeling needed. Their model achieves strong zero-shot results on 3D detection,
-
OpenAI Launches GPT-Realtime-2 for Enhanced Voice Agent Reasoning
By
–
GPT-5 intelligence has officially entered the chat—literally.
— Futurepedia – Learn to Leverage AI (@futurepedia_io) 8 mai 2026
OpenAI just launched GPT-Realtime-2, bringing elite reasoning to voice agents. No more lag, just real-time problem-solving and collaboration.
"Voice agents are now real-time collaborators that can listen, reason,… https://t.co/rSiUD79PdpGPT-5 intelligence has officially entered the chat—literally. OpenAI just launched GPT-Realtime-2, bringing elite reasoning to voice agents. No more lag, just real-time problem-solving and collaboration. "Voice agents are now real-time collaborators that can listen, reason,
-
Thousands of Vibe-Coded Apps Expose Corporate and Personal Data Online
By
–
Thousands of Vibe-Coded Apps Expose Corporate and Personal Data on the Open Web | WIRED https://
share.google/jakCcHeK2SW64e
K2i
… #viben #vibecoding #coding #coder #AI #anthropic #artificialintelligence @AlbertoEMachado @Eli_Krumova @postoff25 @Khulood_Almani @anand_narang @NutritiousMind -
Even image models fail as world models; video is harder
By
–
Even the image models are not good enough world models yet, gpt-image-2 and nano banana make pretty obvious world model like mistakes. is way harder, so no hope of that anytime soon. Maybe with 2-3 OOMs it gets good enough, but that doesn't feel like a reasonable thing to
-
Technical challenges in zero-shot voice synthesis and audio quality
By
–
C’est pas possible d’avoir une voix de bonne qualité (ou de qualité équivalente) en sortie avec un enregistrement GSM compressé. Ça ne va pas ressortir une voix « de mauvaise qualité » : les approches zero-shot vont repartir de données qui ne sont pas entraînées sur des voix GSM
-

Critique of AI voice synthesis output quality
By
–
Et c’est censé prouver quoi ? Je ne comprends pas. Ce que tu démontres, c’est littéralement ce que j’écris dans mon post, non ? Tu as un enregistrement de haute qualité, et pourtant la voix en sortie est très loin d’être de bonne qualité. On est sur une voix monotone, qui plus
