Totally. NVIDIA cosmos is attempting this to pair it with Omniverse as the engine – quality isn’t there yet but they’re pushing.
MULTIMODAL AI
-
xAI Announces Grok Imagine 1.0 for Enhanced Video and Audio Generation
By
–
xAI announced Grok Imagine 1.0 with 10-second videos and improved audio.
— 🚨 AI News | TestingCatalog (@testingcatalog) 2 février 2026
Upgraded video generation is now available on Grok apps and APIs. Generations are also quite different from other top models.
"Cyberpunk robot test" 👀 https://t.co/c79zc7CFhQ pic.twitter.com/G8EbWfR8ofxAI announced Grok Imagine 1.0 with 10-second videos and improved audio. Upgraded video generation is now available on Grok apps and APIs. Generations are also quite different from other top models. "Cyberpunk robot test"
-
User request for expanded multimodal input in Genie AI
By
–
Very doable! Tho I wish genie would just accept video input and multi view images too https://t.co/Z05LOz9lVV
— Bilawal Sidhu (@bilawalsidhu) 2 février 2026Very doable! Tho I wish genie would just accept video input and multi view images too
-

Integrating Generative AI with 3D Scene Graphs and Engines
By
–

Much debate over Genie vs 3D engines. You can have both – the control of 3D scene graphs + the creativity of generative ai. Wrote this in 2024 breaking down the vision. The models are almost there. Now just imagine if Unreal / Unity productized this.
-

Ant Group Releases LingBot-World Open Source Video World Model
By
–
Following Google's Genie 3, Ant Group released an open source alternative: LingBot-World! This video trained world model can also be steered like a game (WASD + camera), with sub-1s latency at 16fps in its real-time setup, and demonstrates coherent rollouts up to 10 minutes,
-

LingBot-Depth: AI-Enhanced Depth Sensing for Robots
By
–
What if your robot or car could see depth more clearly than a top-tier RGB-D camera? Researchers from Ant Group present LingBot-Depth. It treats sensor errors as "masked" clues, using visual context to intelligently fill in and refine incomplete depth maps. It outperforms
-

Explicit 3D Control Meets Generative AI in Virtual Production
By
–
Exactly this. It’s the future of virtual production and gaming. Explicit control of 3D scene graphs + implicit creativity of generative models. Been hammering this since 2023. Only now are models getting good enough to realize it. Y’all should build it!
-
12 Major AI Model Releases Expected in February 2026
By
–
holy sht.. February is going to be insane. 12 major AI drops rumored/expected in the next 30 days: – DeepSeek V4
– ByteDance Doubao 2.0
– Alibaba Qwen 3.5
– Kling 3.0
– Seedance 2.0
– GPT-5.3
– Grok 4.20
– Claude 4.6
– Gemini 3 GA – Apple Gemini-powered Siri
– Meta Avocado
– -

DeepSeek-R1 Simulates Society of Thought for Better Reasoning
By
–
The best AI reasoning might work like a team debate, not a solo monologue. Google & University of Chicago researchers show that top reasoning models like DeepSeek-R1 don't just think longer—they simulate a "society of thought." The model internally hosts diverse expert
-
Moltbook Genie Portals: AI Interface Innovation
By
–
Moltbook curating and spawning genie portals for ppl to jump into?