AI Dynamics

Global AI News Aggregator

About

@shiqi_yang_147

  • WAMs Replace VLAs: Video Models for Advanced Robot Manipulation
    WAMs Replace VLAs: Video Models for Advanced Robot Manipulation

    An elegant and simple pipeline Seonghyeon Ye (@SeonghyeonYe) VLAs (from VLMs) ❌ => WAMs (from Models) ✅ Why WAMs? 1️⃣ World Physics: VLMs know the internet, but Models implicitly model the physical laws essential for manipulation. 2️⃣ The "GPT Direction": VLAs are like BERT (rely heavily on task-specific post-training). WAMs are like GPT (pre-train & prompt), unlocking incredible zero-shot transfer! What I want to see in 2026: 📈 Scaling Laws: We will see much clearer scaling laws for robotics compared to VLAs. 🤝 Human-to-Robot Transfer: Unlocking massive transfer capabilities using video as a shared representation space. 🤖 Zero-Shot Mastery: Moving from short-horizon tasks to long-horizon, dexterous manipulation without task-specific demonstrations. We recently open-sourced the checkpoints, training and inference code. Dive into the research! 👇 📄 Paper: arxiv.org/abs/2602.15922 💻 Code: github.com/dreamzero0/dreamz… 🤗 HF: huggingface.co/GEAR-Dreams/D… — https://nitter.net/SeonghyeonYe/status/2024501978106061056#m

    → View original post on X — @shiqi_yang_147, 2026-02-21 03:30 UTC

  • Vision Redefined: From Representations to Perception-Action Loops

    In my recent blog post, I argue that "vision" is only well-defined as part of perception-action loops, and that the conventional view of computer vision – mapping imagery to intermediate representations (3D, flow, segmentation…) is about to go away. vincentsitzmann.com/blog/bit…

    → View original post on X — @shiqi_yang_147, 2026-02-16 15:31 UTC

  • Chinese Humanoid Robots Rapid Evolution: From Robots to Humans

    You can’t imagine how fast Chinese humanoid robots are evolving. In just one year, they have evolved from robots to "humans". 2025&2026 Chinese Spring Festival Gala

    → View original post on X — @shiqi_yang_147, 2026-02-16 13:22 UTC

  • Impressive AI Video Generator Crushes Everything

    Impressive ——————- This AI video generator crushes everything piped.video/WW_odt7uZTs?si=nSb3…

    → View original post on X — @shiqi_yang_147, 2026-02-11 02:38 UTC

  • Seedance 2.0 Text-to-Video Capabilities Astounding Eerie Mirror Reflection Effect

    🐂🍺 John (@johnAGI168) 😱😱😱 Does Seedance 2.0 model really have no limits to its capabilities? It genuinely gave me quite a shock OMG Seedance 2.0 text-to-video prompt👇 【Style】Pseudo-documentary (Vlog Style), hyper-realism, fixed camera live-action feel, natural lighting, with a touch of suspenseful comedy. 【Duration】15 seconds 【Main character】An ordinary young beautiful woman at the bathroom sink in her home. [00:00-00:06] Shot 1: Daily setup (Normalcy). Scene: In front of an ordinary bathroom mirror. Action: The main character is brushing her teeth in front of the mirror with foam all over her mouth. While brushing, she makes all sorts of silly funny faces at the mirror (raising eyebrows and wrinkling her nose). Key detail: The reflection in the mirror is completely normal at this moment, movements synchronized. [00:06-00:11] Shot 2: The Glitch appears. Action: The main character finishes brushing, lowers her head to spit out the foam, then turns around to leave the bathroom. High-energy moment (core plot twist): Just as the main character's real body has already turned away and left the mirror's frame, the "reflection" in the mirror somehow **doesn't move**! That "reflection" continues to hold the teeth-brushing posture, even smiling wickedly and raising an eyebrow at the camera, staying frozen for a full 2 seconds, before suddenly looking panicked and "fast-forwarding" to catch up with the real body's movements and disappearing. Director's note: Create an extremely realistic "network lag" feeling, as if the reflection has independent consciousness. [00:11-00:15] Shot 3: Comedy callback (The Punchline). Action: The main character, now near the door, seems to sense something wrong and suddenly turns back to look at the mirror. Result: The mirror is now completely normal again, empty, only reflecting the wall opposite. The main character looks utterly confused, scratching her head with a bewildered expression at the camera. The frame freezes on the character's confused face (comedic effect). [Translated from EN to English]

    → View original post on X — @shiqi_yang_147, 2026-02-10 03:37 UTC

  • Seedance 2.0 Generates Shaw Brothers Film Style Short Amazing Results

    OMY Is this AI generated!???? Tractor (@tuolaji2024) Made with seedance 2.0, is AI this powerful now? "Blood Blade Demon Cult", created a short film in Shaw Brothers movie style using the latest Seedance 2.0 model. The entire production relied on just one image of the main characters (shown at the end of the video) and one image of the villain character, with AI handling everything else – the editing, audio-visual synchronization, scene transitions all flow seamlessly, and the results are quite impressive. — https://nitter.net/tuolaji2024/status/2020083327907373103#m [Translated from EN to English]

    → View original post on X — @shiqi_yang_147, 2026-02-07 11:51 UTC

  • Real-time world model generation from text and image prompts

    In our research lab, we are building “real-time dreaming” – the ability to generate fully playable video worlds prompted from any text or image. Our real-time, action conditioned world model (currently running internally at 16fps at 832x480p) is trained on a combination of data, including proprietary Roblox 3D avatar/world interaction data. World models are different from multiplayer engines in that they store state and memory in video latents. Roblox is multiplayer, and we are actively researching optimal ways to simultaneously store state for thousands of players, and keep them in sync with their environment. Our world model leverages database technology which stores all user interactions on Roblox in a vector format that can be used to re-render video and interaction from any camera angle. We see several immediate uses for our Roblox world model. We will use it side-by-side text, image and video prompts as a way to launch auto-generation of immersive worlds. In Roblox Studio, a creator could walk around and use prompts to “paint” a world and then convert it into a 3D representation or direct to Roblox native as a way for many people to play simultaneously. All of this comes alive as we explore the notion of a “Dream Theater” – where one user is dreaming, while others watch and prompt them. 2/4 Community note: The beginning of the video has stolen assets/design from Clair Obscur: Expedition 33 by Sandfall Interactive. The environment is the Flying Waters area and the girl in the video is Maelle. store.steampowered.com/app/1903340/Cl… expedition33.wiki.fextralife.com/Flying+Waters clair-obscur.fandom.com/wiki/Maelle

    → View original post on X — @shiqi_yang_147, 2026-02-05 01:28 UTC

  • Google Labs Project Genie: Pet Explores Infinite Worlds

    amazing Google Labs (@GoogleLabs) Because your pet already thinks the universe revolves around them… why not create one that actually does? With Project Genie, upload a picture of your pet and have them explore infinitely diverse worlds. Learn more: labs.google/projectgenie — https://nitter.net/GoogleLabs/status/2016974664158089425#m

    → View original post on X — @shiqi_yang_147, 2026-01-30 14:24 UTC

  • Project Genie: Advanced World Model Creates Playable Worlds from Text

    Thrilled to launch Project Genie, an experimental prototype of the world's most advanced world model. Create entire playable worlds to explore in real-time just from a simple text prompt – kind of mindblowing really! Available to Ultra subs in the US for now – have fun exploring!

    → View original post on X — @shiqi_yang_147, 2026-01-29 17:23 UTC

  • Low Ego, High Competence: The New Hiring Philosophy for AI Teams

    In a recent chat with a Gemini VP regarding hiring philosophy, one trait he emphasized: the combination of low ego and high competence. We are no longer in an era defined by individual papers or claims of ownership. Success today requires a 'last mile' mindset—a relentless focus on doing whatever work is necessary to deliver world-class models. A team member who pairs high contribution with low ego simplifies and energizes the entire organization. In this hyper-competitive frontier, the delta between contribution and ego has become a key metric for identifying the talent that actually moves the needle.

    → View original post on X — @shiqi_yang_147, 2026-01-24 18:49 UTC