AI Dynamics

Global AI News Aggregator

About

AGENTS

  • CaP Evolution: Agentic Coding and Large Models as Primitives

    Thx Stephen! But quite a bit has changed since 2022…agentic coding is evolving rapidly now and CaP can incorporate large models as primitives. We’re working on extensions and will share updates soon. Stephen James (@stepjamUK) 𝗙𝗿𝗼𝗻𝘁𝗶𝗲𝗿 𝗹𝗮𝗻𝗴𝘂𝗮𝗴𝗲 𝗺𝗼𝗱𝗲𝗹𝘀 𝗰𝗮𝗻 𝗽𝗮𝘀𝘀 𝗹𝗮𝘄 𝗲𝘅𝗮𝗺𝘀. 𝗧𝗵𝗲𝘆 𝗰𝗮𝗻 𝘄𝗿𝗶𝘁𝗲 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻 𝗰𝗼𝗱𝗲. 𝗕𝘂𝘁 𝗮𝘀𝗸 𝘁𝗵𝗲𝗺 𝘁𝗼 𝘄𝗿𝗶𝘁𝗲 𝗮 𝗽𝗿𝗼𝗴𝗿𝗮𝗺 𝘁𝗵𝗮𝘁 𝗰𝗼𝗻𝘁𝗿𝗼𝗹𝘀 𝗮 𝗿𝗲𝗮𝗹 𝗿𝗼𝗯𝗼𝘁, 𝗮𝗻𝗱 𝘁𝗵𝗲𝘆 𝘀𝘁𝗶𝗹𝗹 𝗳𝗮𝗹𝗹 𝘀𝗵𝗼𝗿𝘁 𝗼𝗳 𝗮 𝗵𝘂𝗺𝗮𝗻 𝗲𝘅𝗽𝗲𝗿𝘁. That's the core finding from CaP-X, a new framework from NVIDIA, UC Berkeley, Stanford, and CMU that systematically benchmarks coding agents for robot manipulation. The underlying idea is not new. Code as Policy has been around since 2022/2023, and it is best understood as a modern evolution of Task and Motion Planning – a classical robotics paradigm where engineers manually decompose high-level goals into structured programs combining perception, planning, and control. What has changed is that instead of a human writing that code, a language model does it. It works well when the abstractions are high-level. It degrades significantly when models have to reason at the level human engineers actually work at: raw perception outputs, IK solvers, collision constraints. Here is what the research actually shows: 𝗧𝗵𝗲 𝗮𝗯𝘀𝘁𝗿𝗮𝗰𝘁𝗶𝗼𝗻 𝗴𝗮𝗽 𝗶𝘀 𝗿𝗲𝗮𝗹. Performance drops as you move from high-level primitives to low-level APIs. Not because the models lack intelligence, but because the scaffolding disappears. 𝗠𝘂𝗹𝘁𝗶-𝘁𝘂𝗿𝗻 𝗳𝗲𝗲𝗱𝗯𝗮𝗰𝗸 𝗿𝗲𝗰𝗼𝘃𝗲𝗿𝘀 𝗺𝗼𝘀𝘁 𝗼𝗳 𝘁𝗵𝗮𝘁 𝗹𝗼𝘀𝘀. Multi-turn feedback with execution traces and structured observations dramatically improves performance. Raw images alone actually hurt. 𝗥𝗟 𝗼𝗻 𝗮 𝘀𝗺𝗮𝗹𝗹 𝗺𝗼𝗱𝗲𝗹 𝘁𝗿𝗮𝗻𝘀𝗳𝗲𝗿𝘀 𝘇𝗲𝗿𝗼-𝘀𝗵𝗼𝘁 𝘁𝗼 𝘁𝗵𝗲 𝗿𝗲𝗮𝗹 𝘄𝗼𝗿𝗹𝗱. A 7B model fine-tuned with RL in simulation transfers zero-shot to a real Franka robot by reasoning over structured APIs. The takeaway is simple. The bottleneck is not model size. It is the feedback loop, the abstraction layer, and the system around the model. Credit: @letian_fu, Justin Yu, Karim El-Refai, Ethan Kou, @HaoruXue, @DrJimFan, and the full team across @nvidia, @UCBerkeley, @Stanford, and @CMU_Robotics And of course @AGIBOTofficial for providing the hardware in the attached video! What do you think is holding Code as Policy back from production deployment? Paper link in comments. — https://nitter.net/stepjamUK/status/2041878733531849153#m

    → View original post on X — @ken_goldberg, 2026-04-08 16:41 UTC

  • Motion: AI Video Agent for Motion Design Launch

    introducing Motion, a video agent for tasteful motion design. this launch video was made entirely with Motion. 👇🏽 comment "MOTION" to get early access + free credits. tag @motion_so in any post on your X feed for a surprise. here’s how it works + examples (thread):

    → View original post on X — @scobleizer, 2026-04-08 16:38 UTC

  • AI Agent Management and Human Cognitive Capacity
    AI Agent Management and Human Cognitive Capacity

    Meta just released Muse Spark, the first model from the company's Superintelligence Labs led by Alexandr Wang. Features: natively multimodal, reasoning, tool-use, visual chain of thought, and a "Contemplating mode" that orchestrates multiple agents reasoning in parallel. Some

    → View original post on X — @therundownai

  • Cursor Agents: Control Remote Machines From Your Phone

    You can now run Cursor on any machine and control it from anywhere. Kick off agents from your phone to run on your devbox.

    → View original post on X — @cursor_ai

  • Muse Spark: First Step in AI Scaling Ladder
    Muse Spark: First Step in AI Scaling Ladder

    Muse Spark is the first step on our scaling ladder and the first product of a ground-up overhaul of our AI efforts. It offers competitive performance in multimodal perception, reasoning, health, and agentic tasks. We continue to invest in areas with current performance gaps,

    → View original post on X — @aiatmeta

  • Muse Spark Launches Contemplating Mode for Advanced Reasoning
    Muse Spark Launches Contemplating Mode for Advanced Reasoning

    We're also releasing Contemplating mode, which orchestrates multiple agents that reason in parallel. This allows Muse Spark to compete with the extreme reasoning modes of frontier models such as Gemini Deep Think and GPT Pro. Contemplating will be rolling out gradually in

    → View original post on X — @aiatmeta

  • Meta Introduces Muse Spark, Its New Multimodal Model
    Meta Introduces Muse Spark, Its New Multimodal Model

    Introducing Muse Spark, the first in the Muse family of models developed by Meta Superintelligence Labs. Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration. Muse Spark is available today at meta.ai and the Meta AI app. We're also making it available in private preview via API to select partners, and we hope to open-source future versions of the model. Learn more: go.meta.me/43ea00 [Translated from EN to English]

    → View original post on X — @scobleizer, 2026-04-08 16:05 UTC

  • New Contemplating Mode for Complex Scientific Reasoning Tasks
    New Contemplating Mode for Complex Scientific Reasoning Tasks

    3/ we’re also releasing contemplating mode, which orchestrates multiple agents that reason in parallel designed to handle complex scientific & reasoning queries. in our testing we found it competitive w/ other extreme reasoning models such as Gemini Deep Think & GPT Pro.

    → View original post on X — @alexandr_wang

  • Muse Spark: Meta’s Most Powerful Multimodal Reasoning Model

    2/ muse spark is a natively multimodal reasoning model w/ support for tool-use, visual chain of thought, & multi-agent orchestration. it's the most powerful model that meta has released. through its training process, we saw predictable scaling across pretraining, RL, & test-time

    → View original post on X — @alexandr_wang