Thx Stephen! But quite a bit has changed since 2022โฆagentic coding is evolving rapidly now and CaP can incorporate large models as primitives. Weโre working on extensions and will share updates soon. Stephen James (@stepjamUK) ๐๐ฟ๐ผ๐ป๐๐ถ๐ฒ๐ฟ ๐น๐ฎ๐ป๐ด๐๐ฎ๐ด๐ฒ ๐บ๐ผ๐ฑ๐ฒ๐น๐ ๐ฐ๐ฎ๐ป ๐ฝ๐ฎ๐๐ ๐น๐ฎ๐ ๐ฒ๐
๐ฎ๐บ๐. ๐ง๐ต๐ฒ๐ ๐ฐ๐ฎ๐ป ๐๐ฟ๐ถ๐๐ฒ ๐ฝ๐ฟ๐ผ๐ฑ๐๐ฐ๐๐ถ๐ผ๐ป ๐ฐ๐ผ๐ฑ๐ฒ. ๐๐๐ ๐ฎ๐๐ธ ๐๐ต๐ฒ๐บ ๐๐ผ ๐๐ฟ๐ถ๐๐ฒ ๐ฎ ๐ฝ๐ฟ๐ผ๐ด๐ฟ๐ฎ๐บ ๐๐ต๐ฎ๐ ๐ฐ๐ผ๐ป๐๐ฟ๐ผ๐น๐ ๐ฎ ๐ฟ๐ฒ๐ฎ๐น ๐ฟ๐ผ๐ฏ๐ผ๐, ๐ฎ๐ป๐ฑ ๐๐ต๐ฒ๐ ๐๐๐ถ๐น๐น ๐ณ๐ฎ๐น๐น ๐๐ต๐ผ๐ฟ๐ ๐ผ๐ณ ๐ฎ ๐ต๐๐บ๐ฎ๐ป ๐ฒ๐
๐ฝ๐ฒ๐ฟ๐. That's the core finding from CaP-X, a new framework from NVIDIA, UC Berkeley, Stanford, and CMU that systematically benchmarks coding agents for robot manipulation. The underlying idea is not new. Code as Policy has been around since 2022/2023, and it is best understood as a modern evolution of Task and Motion Planning – a classical robotics paradigm where engineers manually decompose high-level goals into structured programs combining perception, planning, and control. What has changed is that instead of a human writing that code, a language model does it. It works well when the abstractions are high-level. It degrades significantly when models have to reason at the level human engineers actually work at: raw perception outputs, IK solvers, collision constraints. Here is what the research actually shows: ๐ง๐ต๐ฒ ๐ฎ๐ฏ๐๐๐ฟ๐ฎ๐ฐ๐๐ถ๐ผ๐ป ๐ด๐ฎ๐ฝ ๐ถ๐ ๐ฟ๐ฒ๐ฎ๐น. Performance drops as you move from high-level primitives to low-level APIs. Not because the models lack intelligence, but because the scaffolding disappears. ๐ ๐๐น๐๐ถ-๐๐๐ฟ๐ป ๐ณ๐ฒ๐ฒ๐ฑ๐ฏ๐ฎ๐ฐ๐ธ ๐ฟ๐ฒ๐ฐ๐ผ๐๐ฒ๐ฟ๐ ๐บ๐ผ๐๐ ๐ผ๐ณ ๐๐ต๐ฎ๐ ๐น๐ผ๐๐. Multi-turn feedback with execution traces and structured observations dramatically improves performance. Raw images alone actually hurt. ๐ฅ๐ ๐ผ๐ป ๐ฎ ๐๐บ๐ฎ๐น๐น ๐บ๐ผ๐ฑ๐ฒ๐น ๐๐ฟ๐ฎ๐ป๐๐ณ๐ฒ๐ฟ๐ ๐๐ฒ๐ฟ๐ผ-๐๐ต๐ผ๐ ๐๐ผ ๐๐ต๐ฒ ๐ฟ๐ฒ๐ฎ๐น ๐๐ผ๐ฟ๐น๐ฑ. A 7B model fine-tuned with RL in simulation transfers zero-shot to a real Franka robot by reasoning over structured APIs. The takeaway is simple. The bottleneck is not model size. It is the feedback loop, the abstraction layer, and the system around the model. Credit: @letian_fu, Justin Yu, Karim El-Refai, Ethan Kou, @HaoruXue, @DrJimFan, and the full team across @nvidia, @UCBerkeley, @Stanford, and @CMU_Robotics And of course @AGIBOTofficial for providing the hardware in the attached video! What do you think is holding Code as Policy back from production deployment? Paper link in comments. โ https://nitter.net/stepjamUK/status/2041878733531849153#m
โ View original post on X โ @ken_goldberg, 2026-04-08 16:41 UTC