AI Dynamics

Global AI News Aggregator

About

MULTIMODAL AI

  • Vero: Open-Source Vision-Language Model Achieves SOTA Performance

    How do we build a visual AI that truly understands everything from charts to complex science? Researchers at Princeton University present Vero. Vero is a family of fully open-source vision-language models trained with a massive 600K sample dataset (Vero-600K) from 59 diverse datasets, along with a novel reward system. This fully open recipe makes powerful visual reasoning accessible. Vero achieves SOTA performance for open-weight models, improving 3.7-5.5 points across 30 benchmarks on average. It even outperforms Qwen3-VL-8B-Thinking on 23 benchmarks without proprietary thinking data, excelling in spatial reasoning, STEM, chart interpretation, and more.

    → View original post on X — @jiqizhixin

  • RLSD: Self-Distilled Reasoning RL with Token-Level Credit Assignment

    “Self-Distilled RLVR” Most reasoning RL rewards are reliable, but too sparse. Self-Distillation (SD) can fix that with dense token-level signals, but if the teacher sees hidden info, the model can start learning shortcuts it will never have at test time. So this paper, RLSD, let RL decide whether an answer was good or bad, and let self-distillation decide which tokens deserve more credit. And instead of using a teacher to tell the model what to imitate, they use it to do token-level credit assignment, which gives denser learning than vanilla RLVR, without the instability and leakage of naive self-distillation. Empirically, RLSD stays stable while on-policy SD degrades, and beats GRPO-style baselines on multimodal reasoning.

    → View original post on X — @askalphaxiv

  • OpenClaw 2026.4.7 Brings Native LLM Memory Wiki Support

    Proud to bring fully native @karpathy's LLM wiki support including backfilling, native @obsdmd, and intergration with /dreams. 🧠 Memory features seem to be the next big unlock for agentic systems. OpenClaw🦞 (@openclaw) OpenClaw 2026.4.7 🦞 🔮 openclaw infer 🎬 music + video editing 💾 session branch/restore 🔗 webhook-driven TaskFlows 🤖 Arcee, Gemma 4, Ollama vision 🧠 memory-wiki: persistent knowledge, not just vibes Because “trust me bro” is not a knowledge system. github.com/openclaw/openclaw… — https://nitter.net/openclaw/status/2041714270212108657#m

    → View original post on X — @scobleizer, 2026-04-08 06:09 UTC

  • Egocentric-1M: Largest Egocentric Video Dataset for Physical AI

    introducing Egocentric-1M. the largest egocentric video dataset in the world, and our next step in building the internet for physical AI. Eddy Xu (@eddybuild) today, we’re open sourcing the largest egocentric dataset in history. – 10,000 hours – 2,153 factory workers – 1,080,000,000 frames the era of data scaling in robotics is here. (thread) — https://nitter.net/eddybuild/status/1987951619804414416#m

    → View original post on X — @scobleizer, 2026-04-08 05:34 UTC

  • OpenClaw 2026.4.7 Release: Inference, Editing, Memory Wiki

    Second 🚢 of the day. OpenClaw🦞 (@openclaw) OpenClaw 2026.4.7 🦞 🔮 openclaw infer 🎬 music + video editing 💾 session branch/restore 🔗 webhook-driven TaskFlows 🤖 Arcee, Gemma 4, Ollama vision 🧠 memory-wiki: persistent knowledge, not just vibes Because “trust me bro” is not a knowledge system. github.com/openclaw/openclaw… — https://nitter.net/openclaw/status/2041714270212108657#m

    → View original post on X — @ceobillionaire, 2026-04-08 03:12 UTC

  • OpenClaw 2026.4.7 Release: AI Inference, Media Editing, Knowledge System

    OpenClaw 2026.4.7 🦞 🔮 openclaw infer 🎬 music + video editing 💾 session branch/restore 🔗 webhook-driven TaskFlows 🤖 Arcee, Gemma 4, Ollama vision 🧠 memory-wiki: persistent knowledge, not just vibes Because “trust me bro” is not a knowledge system. github.com/openclaw/openclaw…

    → View original post on X — @ceobillionaire, 2026-04-08 03:06 UTC

  • Psi-Zero: Open Foundation Model for Humanoid Robot Learning

    What if we could teach humanoid robots intricate skills more efficiently than ever before? The USC Physical Superintelligence (PSI) Lab, NVIDIA, and WorldEngine introduce Ψ0 (Psi-Zero). Their new open foundation model rethinks how humanoids learn complex tasks by decoupling the learning process: it first acquires general visual-action understanding from human videos, then masters precise robot control using high-quality humanoid data. Ψ0 sets a new standard for universal humanoid loco-manipulation, achieving over 40% higher success rates across multiple complex tasks while using more than 10 times less training data than prior state-of-the-art approaches. Ψ0: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation Paper: arxiv.org/abs/2603.12263 Project: psi-lab.ai/Psi0/ Code: github.com/physical-superint… Our report: mp.weixin.qq.com/s/yvkG5ZcO1… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin

  • VR and Latent Space Replace Psychedelics

    there won't be a need for psychedelics anymore just vr goggles and latent space Benjamin Bardou (@benjaminbardou) Vertigo of Latent Space Montage marked the decisive invention of cinema: a mode of associating images capable of producing new relations. Yet this potential has remained largely unexplored, as cinema has most often reproduced narrative forms inherited from the novel. With artificial imagination, another regime of images emerges. Images are no longer simply arranged; they transform into one another within a continuous space. It then becomes possible to work not with sequences, but with transitions, passages, thresholds. Vertigo is approached here as a latent structure, a field of forms in circulation. The film becomes a set of persistent forms that reactivate, deform, and recombine in contact with other images. What is at stake here is less a reinterpretation of the film than an attempt to approach thought in action: its movement, its bifurcations, its reminiscences. In The Flow, this research unfolds on another scale. It seeks to follow the flow of consciousness, not as narrative, but as a continuous dynamic in which films, memories, and history intermingle and circulate. This work around Vertigo constitutes a variation: a way of exploring how a film can dissolve into this flow and become one of the sites from which thought begins to move. — https://nitter.net/benjaminbardou/status/2041519539234508968#m

    → View original post on X — @scobleizer, 2026-04-08 02:15 UTC

  • Robots and Humans in Shared Virtual Spaces: Beyond World Models

    Interesting post by @peteflorence who is a robotics pioneer. He triggered me by positioning against people who talk about technology like world models. But he is right. Where we are going with robotics is the interesting part. It is clear to me that both humans and robots will work and play together in a shared Holodeck. The nerds call them SLAM maps but they basically turn your house, or wherever, into grids of tiny cubes with a virtual display on each side. Like 1 mm cubes. Voxels. I can see these when I turn on my Apple Vision Pro. And the computer can change the display into anything else. I turn the knob on my headset and it changes my kitchen to Yosemite National Park. It can change the clothing on your robot too. Which gets to the point. A decade from now we aren’t going to care what is running the robot. World Model or whatever AGI-driven technology soon will come. Just that it does stuff for us and with us. My expectations are that it will do anything you can think of doing. Not there yet. Pete Florence (@peteflorence) x.com/i/article/204137153848… — https://nitter.net/peteflorence/status/2041529286562402804#m

    → View original post on X — @scobleizer, 2026-04-08 00:16 UTC

  • AI Image Generation Capabilities: Opossum on E-Scooter Example
    AI Image Generation Capabilities: Opossum on E-Scooter Example

    And for the "they're training on your pelican now!" skeptics, here's what I got for NORTH VIRGINIA OPOSSUM ON AN E-SCOOTER Comments on that one include /* Earring sparkle */, ,

    → View original post on X — @simonw