AI Dynamics

Global AI News Aggregator

About

CODE

  • Move to Open Source Models from Hugging Face

    Time to move to open or local models from Hugging Face! All instructions are here: huggingface.co/blog/liberate… Boris Cherny (@bcherny) Starting tomorrow at 12pm PT, Claude subscriptions will no longer cover usage on third-party tools like OpenClaw. You can still use these tools with your Claude login via extra usage bundles (now available at a discount), or with a Claude API key. — https://nitter.net/bcherny/status/2040206440556826908#m

    → View original post on X — @huggingface, 2026-04-04 00:06 UTC

  • 7 Essential Python Itertools for Feature Engineering
    7 Essential Python Itertools for Feature Engineering

    7 Essential Python Itertools for Feature Engineering https://machinelearningmastery.com/7-essential-python-itertools-for-feature-engineering/?utm_source=dlvr.it&utm_medium=twitter [Translated from EN to English]

    → View original post on X — @craigbrownphd, 2026-04-04 00:02 UTC

  • Open Source Contributions Improve Prompt Cache Efficiency

    We're big fans of open source. I actually just put up a few PRs to improve prompt cache efficiency for OpenClaw specifically. This is more about engineering constraints. Our systems are highly optimized for one kind of workload, and to serve as many people as possible with the

    → View original post on X — @bcherny

  • Coding Agents Require Significant Cognitive Load and Burnout Risk

    Just noticed this has had 1.1m views now, which explains why I starting to see some less informed reactions to it starting to crop up now it's broken out of purely tech Twitter Lenny Rachitsky (@lennysan) "Using coding agents well is taking every inch of my 25 years of experience as a software engineer, and it is mentally exhausting. I can fire up four agents in parallel and have them work on four different problems, and by 11am I am wiped out for the day. There is a limit on human cognition. Even if you're not reviewing everything they're doing, how much you can hold in your head at one time. There's a sort of personal skill that we have to learn, which is finding our new limits. What is a responsible way for us to not burn out, and for us to use the time that we have?" @simonw — https://nitter.net/lennysan/status/2039845666680176703#m

    → View original post on X — @simonw, 2026-04-03 23:41 UTC

  • Working on Claude Code improvements, use latest version

    Actively working on making it better. Make sure you're using the latest Claude Code version

    → View original post on X — @bcherny

  • Closing the 100,000-Year Robot Data Gap with Code-as-Policy
    Closing the 100,000-Year Robot Data Gap with Code-as-Policy

    Thank your for this excellent summary Junfan! Junfan Zhu 朱俊帆 (@junfanzhu98) How to Close the 100,000-Year Robot “Data Gap” — @Ken_Goldberg (@UCBerkeley) Goldberg’s core claim: end-to-end Vision-Language-Action (VLA) models aren’t delivering. They’re opaque, hard to debug, and fragile under distribution shift. On LIBERO / LIBERO-PRO, models reach ~100% in-distribution, but tiny pose perturbations collapse success to ~17% or 0% (even π₀). This is systematic overfitting, not generalization. Code-as-Policy (CaP) reframes control: LLMs generate executable programs that call structured primitives (perception, 6D pose, motion planning, grasping). Generalization shifts from weights → code. Benefits: interpretable, verifiable, training-free at inference, debuggable. Open question: reliability. CaP-X (arXiv 2603.22435) introduces a full evaluation stack: 🔷 CaP-Gym: unified REPL over RoboSuite + LIBERO-PRO + BEHAVIOR (tabletop → mobile/bimanual, sim→real) 🔷 CaP-Bench: multi-level abstraction tests 🔷 CaP-Agent0: training-free agent (visual differencing, skill library, parallel queries) 🔷 CaP-RL: verifiable reward RL in Python sandbox Results: under perturbations that break VLAs, CaP-X hits ~96% success with strong pose invariance (extreme corners, lighting, object swaps). LIBERO-PRO (50 trials/task): many 100%, lowest ~76–78%. Failures are mostly semantic (label ambiguity), not control. Grasping/planning ≈ solved. CaP-X 2.0 pushes agentic coding: prompt restructuring, failure-analysis primitives, human-in-loop, reusable cloud skill cache. Test-time loop (no retraining): generate → compile → execute → perturb → diagnose → patch. Extensions include Rust backends (reliability) and Graph-as-Policy (GaP) for node-level verification. Core thesis (GOFE + CaP hybrid): pure VLA scaling cannot close the 100,000-year data gap (robot ≈10K hrs vs LLM ≈1.2B hrs). Robotics needs a Good Old-Fashioned Engineering (GOFE) skeleton: modular pipelines, PID (kp, kv), feedforward (e.g., virtual gravity) — inherently pose-invariant. → Build GOFE backbone + CaP brain. Deploy now, collect real data, spin a flywheel to improve modules and future VLAs. Hot 🔥 takes: 🔷 “Robot generalists should get off their high horse.” VLA-only is dogmatic. 🔷 VLAs may win eventually, but near-term progress requires hybrids. 🔷 Reliability (→99.9%) is the real bottleneck, not demos. Comparison 🔷 GOFE: no generality, but available, interpretable, reliable 🔷 VLA: promised generality, but opaque, brittle, not ready 🔷 CaP: generality + available + interpretable; reliability improves via hybrid + iteration Timeline: near-term: structured tasks (declutter, laundry, delivery). ~5 years: major home logistics gains. Full humanoid generalists: far off. Strategy: specialists first + reliability to 99.9%. Bottom line: don’t wait for end-to-end intelligence. Turn LLMs into super-programmers over a GOFE substrate, deploy hybrids, iterate with real data, and asymptotically approach VLA—without burning out the field. 👉🏻More pics: linkedin.com/posts/junfan-zh… — https://nitter.net/junfanzhu98/status/2039953079706247169#m

    → View original post on X — @ken_goldberg, 2026-04-03 23:15 UTC

  • Welcome to OpenAI Team Building Codex

    welcome to the team!! Kath Korevec (@simpsoka) Can’t wait to join the team at @openai building codex. Would love to hear what you love about it or want changed. We’re moving fast. DMs open. — https://nitter.net/simpsoka/status/2040144952550969516#m

    → View original post on X — @gdb, 2026-04-03 23:09 UTC

  • Jeff Dean’s First PR in Years Contributed to Hugging Face Transformers

    First public PR in years from the legend @JeffDean was to @huggingface transformers. Any other proud moment for us and our community! Omar Sanseviero (@osanseviero) New Jeff Dean fact: the transformers PR for Gemma 4 had 14 authors and Jeff was one of them — https://nitter.net/osanseviero/status/2040177838817476879#m

    → View original post on X — @huggingface, 2026-04-03 22:53 UTC

  • Claude Code Feature Limitations: Interactive Prompting Gap

    The feature set is the problem – the model and harness are great, but in Claude Code for web I can do things like prompt the model while it's working

    → View original post on X — @simonw

  • Devin AI Achieves Agentic Self-Improvement with One-Shot Implementation
    Devin AI Achieves Agentic Self-Improvement with One-Shot Implementation

    We have achieved agentic self improvement – i can just copy paste blogposts and tweets into @devinai and it oneshots the complete implementation wasnt actually sure this was gonna work, jaw dropped when it did. this is very out of distribution of the underlying @GoogleDeepMind Gemini Flash Lite model but it Just Worked. Malte Ubl (@cramforce) Mintlify assistant is powered by just-bash with a custom filesystem — https://nitter.net/cramforce/status/2039841201474695333#m

    → View original post on X — @swyx, 2026-04-03 21:34 UTC