AI Dynamics

Global AI News Aggregator

About

CODE

  • LangSmith: Tracing and Evals for Agent Optimization

    LangSmith 🤝 San Francisco You don't know what your agents will do until you actually run them. What works in demos can break in the real world. Without tracing and evals, you're just guessing at why. Track what your agent actually does. Optimize and fix your agents. Then measure whether your fixes work. That loop is how agents get better, and LangSmith is built to power that workflow.

    → View original post on X — @langchain

  • Apple blocks Anything app despite serving millions of vibe coders
    Apple blocks Anything app despite serving millions of vibe coders

    .@Apple will literally approve 50 identical 'Chat with AI Waifu' scam apps before noon.. .. but they block a cool vibe-coding tool that actually teaches kids how to code. Make it make sense. Anything (@anything) Guideline 2.5.2 – Gatekeeping – Vibes denied we haven't talked about this publicly for months we tried to resolve it privately with emails, calls, appeals, and four technical rewrites to comply with whatever Apple wanted here's our truth, unfiltered on March 26th, Apple removed Anything from the App Store then they brought us back now they removed us again and I think it's time to say something, because this isn't really about us. It's about who gets to build software, and who gets to decide for most of the history of computing, making an app required years of specialized training. You either knew how to code or you didn't, and if you didn't, your idea stayed in your head forever. that barrier is falling right now. Millions of people are discovering they can describe what they want and get a working app they call themselves vibe coders and they are the most exciting audience in technology they're building things nobody else would have built because nobody else had their problems a firefighter in Northern California used Anything to build an emergency incident response app he never wrote a line of code. Did hundreds of iterations, testing each one on his iPad through our mobile preview app got it into the App Store. Now he's selling it to fire departments across the state. it would have cost him over a hundred thousand dollars to hire engineers He spent a few hundred bucks. That guy is why we exist. Not the technology. Him. And the millions of people like him. our mobile app did one thing for people like him it let them preview what they were building with Anything on their own phone. GPS, camera, notifications, things you can only test on a real device with native code They'd iterate, try it, tweak it, try again. When they were happy, they'd submit to the App Store through the normal process Apple reviewed it like any other app. Our mobile app got approved last year. We didn't hear a word of concern. then in December, they started blocking our updates, citing the infamous Guideline 2.5.2 the rule designed to prevent malicious apps from downloading code to change their behavior after review We understood the concern, even if we disagree it applies to us. We tried to fix it. Four different technical approaches, each one specifically designed to address what they told us. Each one rejected. we didn't go public we didn't tweet we kept trying then they pulled us from the App Store. We still didn't say anything. We worked with them, got reinstated, believed we'd found a path forward Then they pulled us again. at some point silence stops being patience and starts being complicity. We have builders who depend on us. They deserve to know what's happening and why. Guideline 2.5.2 is a good rule. apps shouldn't be able to pass review and then become something else. But that's not us. We help people preview their own work on their own device Expo Go has done the exact same thing for professional developers for years and is on the App Store right now, today! the only difference is our users aren't professional developers they're the firefighter they're the teacher building a classroom app they're the person who discovered last week that they could build software at all that's who Apple is locking out. Not us. Them. and here's what I need Apple to understand these people are the future of the App Store. Not a sideshow. The future. The number of people who can build apps is about to go from millions to hundreds of millions to eventually everyone the platforms and tools that serve those people will determine where they build every vibe coder who ships through Anything is a new developer in Apple's ecosystem who didn't exist a year ago They want to build web apps, Android apps, and yes iOS apps we help them add in-app purchases. We help them make their apps secure and scale. We catch rejection issues early. We are a feeder system for the App Store The safety argument is hollow. Preview apps only run on the builder's own device. They're sandboxed in the Anything mobile app. Want anyone else to use it? You still submit to the App Store. Apple still reviews every line. We're not bypassing review. We're a dress rehearsal for it. but none of that matters when a reviewer sees "downloads executable code" on a checklist and reaches for reject without asking what the code is, how it actually works, or who it's for. we're not waiting we launched text-to-app. Text us and we'll build your iOS app in the cloud We're shipping a desktop companion for on-device previews next. We'll find a way to serve our builders We always do. but I'm done being quiet about why we have to the people we serve, the ones crazy enough to start their own thing, building apps for their fire departments and their classrooms and their small businesses they deserve to test what they're making on the device it's made for that's not a loophole that's how building works – Apple can be the platform where the next hundred million builders get started – or they can keep banning the tools those people depend on and watch it happen somewhere else we all know which one the firefighter will choose — https://nitter.net/anything/status/2041599393237774507#m

    → View original post on X — @datachaz, 2026-04-07 22:44 UTC

  • Rocket 1.0 Launches: Persistent Context Between Sessions for Product Development

    "NOTHING resets between sessions." Anyone who's ever copy-pasted a 3,000-word prompt into ChatGPT for the 50th time just to remind it what app you're building knows exactly how MASSIVE this is! Huge congrats to the @rocketdotnew team 🤘 Vishal Virani (@Vishalvirani91) Rocket 1.0 is live. This is our first major step toward Vibe Solutioning. Vibe coding solved how to build. It never solved what to build, or why. That's the harder problem and the one where most products actually fail. @rocketdotnew connects the thinking and the building in one platform. Solve your hardest business question. Build from what you solved. Watch your competition while you work. Everything shares one context. Nothing resets between sessions. The video and blog explain it better than I can here. — https://nitter.net/Vishalvirani91/status/2041546557342855363#m

    → View original post on X — @datachaz, 2026-04-07 22:02 UTC

  • Vibe Coding Explained for Beginners and App Builders

    Curious about vibe coding? Or are you already shipping apps and just want an easier way to explain your new favorite hobby to your friends, parents, grandparents, etc.? Either way, this video is for you

    → View original post on X — @googleai

  • Rocket shifts focus from how to build to what and why

    Vibe coding has been solving the wrong problem this whole time. Everyone optimized for how to build. Rocket is the first tool asking what to build and why, that framing shift changes everything.

    → View original post on X — @aihighlight

  • Karpathy’s Self-Improving AI Knowledge Base with Obsidian
    Karpathy’s Self-Improving AI Knowledge Base with Obsidian

    ICYMI here's more info about Andrej’s new method nitter.net/DataChaz/status/203996… Charly Wargnier (@DataChaz) 🚨 Karpathy’s new set-up is the ultimate self-improving second brain, and it takes zero manual editing 🤯 It acts as a living AI knowledge base that actually heals itself. Let me break it down. Instead of relying on complex RAG, the LLM pulls raw research directly into an @Obsidian Markdown wiki. It completely takes over: ✦ Index creation ✦ System linting ✦ Native Q&A routing The core process is beautifully simple: → You dump raw sources into a folder → The LLM auto-compiles an indexed .md wiki → You ask complex questions → It generates outputs (Marp slides, matplotlib plots) and files them back in The big-picture implication of this is just wild. When agents maintain their own memory layer, they don’t need massive, expensive context limits. They really just need two things: → Clean file organization → The ability to query their own indexes Forget stuffing everything into one giant prompt. This approach is way cheaper, highly scalable… and 100% inspectable! — https://nitter.net/DataChaz/status/2039963758790156555#m

    → View original post on X — @datachaz, 2026-04-07 21:15 UTC

  • Karpathy’s Autonomous Obsidian Wiki System Replaces Traditional RAG
    Karpathy’s Autonomous Obsidian Wiki System Replaces Traditional RAG

    🚨 @karpathy literally ditched traditional RAG for an autonomous Obsidian file system. Instead of writing code, he dumps raw AI research into a local folder and lets an LLM convert it into an interconnected markdown wiki. He rarely edits the text manually. By relying purely on dynamically updated index files, the system navigates the exact context it needs natively without relying on flawed vector embeddings. Because the LLM fully understands the file structure, it executes advanced autonomous workflows: → Operates a custom vibe-coded local search engine → Renders complex charts and formatted markdown slides → Continuously compounds a 400,000-word knowledge base The most fascinating mechanic is the self-healing loop. He triggers background health checks where the LLM natively spots structural gaps, scrapes the internet for missing data, and cleans the articles perfectly. This feels the absolute blueprint for managing complex technical data 🔥 btw, he also plans to fine-tune a local model directly on the wiki so the research is baked into the neural weights rather than relying on limited context windows 👀

    → View original post on X — @datachaz, 2026-04-07 21:10 UTC

  • Claude Mythos: Ten Trillion Parameter Model Deployed for Cybersecurity

    Claude Mythos. Ten trillion parameters: the first model in this weight class. Estimated training cost: ten billion dollars. On the hardest coding test in the industry (SWE bench) it scores 94%. It found a security flaw in a system that had been running for 27 years, one that every human engineer and every automated check had missed. It found another bug that had survived five million test runs over 16 years. (It did so overnight.) It is so capable in cybersecurity that Anthropic will not release it to the public, instead it is launching Project Glasswing along with 100m in compute credits to help secure software. Only twelve partners currently have access: Amazon, Cisco, Apple, Google, Microsoft, NVIDIA, JPMorgan Chase, Crowdstrike, Palo Alto, AWS, The Linux Foundation, Broadcom. (I'm sure the Pentagon is on the line?) This is not a product launch: it is a controlled deployment of a system too powerful to distribute freely. Tell me this isn't (very expensive) AGI? Anthropic (@AnthropicAI) Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, Claude Mythos Preview, which can find software vulnerabilities better than all but the most skilled humans. anthropic.com/glasswing — https://nitter.net/AnthropicAI/status/2041578392852517128#m

    → View original post on X — @ceobillionaire, 2026-04-07 21:04 UTC

  • Claude Mythos Achieves AGI: Perfect Hardware Design Generation

    Just got access to Claude Mythos… & ughhhhhhhhh this is AGI. It was the first time a model one shotted a 10/25G Ethernet MAC/PCS, it even knew to select the right line rate and data width for lower latency. This alone is something that would take a really skilled digital designer 3-6 months if they had experience in the past to pull off… But it didn’t just do that I then said to make the MAC fully cut through and only forward certain IP addresses within a range downstream it one shotted it instantly also which blew me away… Then finally I thought ok let me trip it up so I said now do 50G MAC and it knew without me telling it to add another GT transceiver and it even added alignment markers and FEC to it correctly. 💀💀💀 It’s passing all the tests I have so I’m going to flash the board and see if it actually works on hardware now…

    → View original post on X — @ceobillionaire, 2026-04-07 21:00 UTC

  • BadClaude tool released on GitHub with npm installation

    still wanna try the whip? it's here 👿 github.com/GitFrog1111/badcl… npm install -g badclaude badclaude

    → View original post on X — @datachaz, 2026-04-07 20:49 UTC