AI Dynamics

Global AI News Aggregator

About

CODE

  • Hamiltonian Monte Carlo: Physics-Based Probabilistic Sampling

    Hamiltonian Monte Carlo: probability as physics. Endow particles with momentum, then let Hamilton’s equations (dq/dt = ∂H/∂p, dp/dt = −∂H/∂q) carve reversible, volume-preserving trajectories through phase space.

    → View original post on X — @ceobillionaire

  • Delayed Message Feature Request for Background Job Notifications

    Codex feature request: 'delayed message'; I often have long running jobs in the background and I want to send a 'how is it going?' message in, say 20 mins or an hour. I don't want to create an automation for this, just a delay is enough. Would be a nice quality of life feature

    → View original post on X — @petergostev

  • Anthropic Launches Managed Agents for Unpredictable Programs

    New on the Engineering Blog: Building Managed Agents—our hosted service for long-running agents—meant solving an old problem in computing: how to design a system for "programs as yet unthought of." Read more: anthropic.com/engineering/managed-agents [Translated from EN to English]

    → View original post on X — @anthropicai, 2026-04-08 17:20 UTC

  • Self-Improving Agents: Systems Engineering and Evaluation Infrastructure

    Self-improving agents isn’t a single algorithm – it’s a systems engineering problem involving: – eval data curation + maintenance – experiment design to battle overfitting – an update algorithm – human review during the process & especially before prod we share practical learnings + a local research scaffold to autonomously hill-climb harness centered around evals our goal is to give everyone the tooling and infra to measure and iteratively their improve agents. Evals are training data for agents which fuels this loop let's build the future of well-designed, self-improving systems 🚀 Viv (@Vtrivedy10) x.com/i/article/204172946391… — https://nitter.net/Vtrivedy10/status/2041927488918413589#m

    → View original post on X — @langchain

  • Meta’s Muse Spark Converts Images to Code with Asset Extraction

    Ok this is actually pretty impressive and I truly didn't see any model doing this before or being able to do it to this extent. When I asked Muse Spark from Meta to convert this image into code, it cut out the assets from the screens so it could use them correctly!

    → View original post on X — @alexandr_wang, 2026-04-08 17:08 UTC

  • Open Models Found FreeBSD Zero-Day Vulnerability Across Tasks

    New post: We tested the Mythos showcase vulnerabilities with open models. They recovered similar scoped analysis! 8/8 models found the flagship FreeBSD zero-day, including a 3B model. Rankings reshuffle completely across tasks => the AI cybersecurity frontier is super jagged!

    → View original post on X — @clementdelangue

  • CaP Evolution: Agentic Coding and Large Models as Primitives

    Thx Stephen! But quite a bit has changed since 2022…agentic coding is evolving rapidly now and CaP can incorporate large models as primitives. We’re working on extensions and will share updates soon. Stephen James (@stepjamUK) 𝗙𝗿𝗼𝗻𝘁𝗶𝗲𝗿 𝗹𝗮𝗻𝗴𝘂𝗮𝗴𝗲 𝗺𝗼𝗱𝗲𝗹𝘀 𝗰𝗮𝗻 𝗽𝗮𝘀𝘀 𝗹𝗮𝘄 𝗲𝘅𝗮𝗺𝘀. 𝗧𝗵𝗲𝘆 𝗰𝗮𝗻 𝘄𝗿𝗶𝘁𝗲 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻 𝗰𝗼𝗱𝗲. 𝗕𝘂𝘁 𝗮𝘀𝗸 𝘁𝗵𝗲𝗺 𝘁𝗼 𝘄𝗿𝗶𝘁𝗲 𝗮 𝗽𝗿𝗼𝗴𝗿𝗮𝗺 𝘁𝗵𝗮𝘁 𝗰𝗼𝗻𝘁𝗿𝗼𝗹𝘀 𝗮 𝗿𝗲𝗮𝗹 𝗿𝗼𝗯𝗼𝘁, 𝗮𝗻𝗱 𝘁𝗵𝗲𝘆 𝘀𝘁𝗶𝗹𝗹 𝗳𝗮𝗹𝗹 𝘀𝗵𝗼𝗿𝘁 𝗼𝗳 𝗮 𝗵𝘂𝗺𝗮𝗻 𝗲𝘅𝗽𝗲𝗿𝘁. That's the core finding from CaP-X, a new framework from NVIDIA, UC Berkeley, Stanford, and CMU that systematically benchmarks coding agents for robot manipulation. The underlying idea is not new. Code as Policy has been around since 2022/2023, and it is best understood as a modern evolution of Task and Motion Planning – a classical robotics paradigm where engineers manually decompose high-level goals into structured programs combining perception, planning, and control. What has changed is that instead of a human writing that code, a language model does it. It works well when the abstractions are high-level. It degrades significantly when models have to reason at the level human engineers actually work at: raw perception outputs, IK solvers, collision constraints. Here is what the research actually shows: 𝗧𝗵𝗲 𝗮𝗯𝘀𝘁𝗿𝗮𝗰𝘁𝗶𝗼𝗻 𝗴𝗮𝗽 𝗶𝘀 𝗿𝗲𝗮𝗹. Performance drops as you move from high-level primitives to low-level APIs. Not because the models lack intelligence, but because the scaffolding disappears. 𝗠𝘂𝗹𝘁𝗶-𝘁𝘂𝗿𝗻 𝗳𝗲𝗲𝗱𝗯𝗮𝗰𝗸 𝗿𝗲𝗰𝗼𝘃𝗲𝗿𝘀 𝗺𝗼𝘀𝘁 𝗼𝗳 𝘁𝗵𝗮𝘁 𝗹𝗼𝘀𝘀. Multi-turn feedback with execution traces and structured observations dramatically improves performance. Raw images alone actually hurt. 𝗥𝗟 𝗼𝗻 𝗮 𝘀𝗺𝗮𝗹𝗹 𝗺𝗼𝗱𝗲𝗹 𝘁𝗿𝗮𝗻𝘀𝗳𝗲𝗿𝘀 𝘇𝗲𝗿𝗼-𝘀𝗵𝗼𝘁 𝘁𝗼 𝘁𝗵𝗲 𝗿𝗲𝗮𝗹 𝘄𝗼𝗿𝗹𝗱. A 7B model fine-tuned with RL in simulation transfers zero-shot to a real Franka robot by reasoning over structured APIs. The takeaway is simple. The bottleneck is not model size. It is the feedback loop, the abstraction layer, and the system around the model. Credit: @letian_fu, Justin Yu, Karim El-Refai, Ethan Kou, @HaoruXue, @DrJimFan, and the full team across @nvidia, @UCBerkeley, @Stanford, and @CMU_Robotics And of course @AGIBOTofficial for providing the hardware in the attached video! What do you think is holding Code as Policy back from production deployment? Paper link in comments. — https://nitter.net/stepjamUK/status/2041878733531849153#m

    → View original post on X — @ken_goldberg, 2026-04-08 16:41 UTC

  • 5 AI Model Architectures Every Engineer Should Know
    5 AI Model Architectures Every Engineer Should Know

    5 #AI Model Architectures Every AI Engineer Should Know by Arham Islam @Marktechpost Learn more: bit.ly/4s5g1pA #LLM #ArtificialIntelligence #GenerativeAI #ML #MachineLearning

    → View original post on X — @ronald_vanloon, 2026-04-08 16:25 UTC

  • API Release Coming Soon to Power Applications

    we intend to release the API soon! we look forward to it powering some claws out there 🙂

    → View original post on X — @alexandr_wang

  • AI Adoption in Enterprise: Data and Sectoral Trends
    AI Adoption in Enterprise: Data and Sectoral Trends

    Enterprises are using AI today for coding, legal, support, healthcare, and more. @kimberlywtan's must-read deep dive compiles hard data on where AI has the most enterprise adoption – and the industries AI is coming for next: a16z.news/p/ai-adoption-by-t… [Image] Kimberly Tan (@kimberlywtan) x.com/i/article/204175770374… — https://nitter.net/kimberlywtan/status/2041896368877531158#m [Translated from EN to English]

    → View original post on X — @scobleizer, 2026-04-08 16:16 UTC