AI Dynamics

Global AI News Aggregator

About

LLMS

  • GLM-5.1 weights open-source, beats top AI models on SWE-Bench
    GLM-5.1 weights open-source, beats top AI models on SWE-Bench

    INCREDIBLE GLM-5.1 weights are now opensource > i’ve had early access to the weights for the past few days
    > and yeah… this one matters a lot benchmarks? > SWE-Bench Pro: 58.4
    > beats Opus 4.6 (57.3)
    > beats GPT-5.4 (57.7)
    > beats Gemini 3.1 Pro (54.2) let that sink in

    → View original post on X — @theahmadosman

  • Anthropic’s Internal Mythos Project Since February 2024
    Anthropic’s Internal Mythos Project Since February 2024

    Explains the crazy rate of shipping Lisan al Gaib (@scaling01) ANTHROPIC HAD MYTHOS INTERNALLY SINCE FEB 24 — https://nitter.net/scaling01/status/2041587896541499543#m

    → View original post on X — @ceobillionaire, 2026-04-07 23:08 UTC

  • Claude Mythos AI breakthrough obliterates all benchmarks
    Claude Mythos AI breakthrough obliterates all benchmarks

    This is insane. Truly the end times. Deedy (@deedydas) Claude Mythos just obliterated every single benchmark in AI. I can't believe what I'm reading. — https://nitter.net/deedydas/status/2041605983659860115#m

    → View original post on X — @ceobillionaire, 2026-04-07 21:54 UTC

  • Claude Mythos: Anthropic’s Token Efficiency Breakthrough and IPO Prospects
    Claude Mythos: Anthropic’s Token Efficiency Breakthrough and IPO Prospects

    Claude Mythos is not only a big leap in performance, it's also about 5x token efficient in BrowseComp. I don't know what Anthropic is doing. But they manage to surprise me every single time. The IPO is getting closer. They have an ARR OpenAI outrun with $30 billion in revenue. OpenAI is under pressure. The next release has to be a huge hit because the market is evaluating the future. The pressure couldn't be greater. OpenAI has to prove its own "Mythos."

    → View original post on X — @kimmonismus, 2026-04-07 21:34 UTC

  • Pelican GLM-5.1 Impresses with Animation Capabilities

    I'm a big fan of the pelican GLM-5.1 drew me today, it even animated it!

    → View original post on X — @simonw

  • Karpathy’s Self-Improving AI Knowledge Base with Obsidian
    Karpathy’s Self-Improving AI Knowledge Base with Obsidian

    ICYMI here's more info about Andrej’s new method nitter.net/DataChaz/status/203996… Charly Wargnier (@DataChaz) 🚨 Karpathy’s new set-up is the ultimate self-improving second brain, and it takes zero manual editing 🤯 It acts as a living AI knowledge base that actually heals itself. Let me break it down. Instead of relying on complex RAG, the LLM pulls raw research directly into an @Obsidian Markdown wiki. It completely takes over: ✦ Index creation ✦ System linting ✦ Native Q&A routing The core process is beautifully simple: → You dump raw sources into a folder → The LLM auto-compiles an indexed .md wiki → You ask complex questions → It generates outputs (Marp slides, matplotlib plots) and files them back in The big-picture implication of this is just wild. When agents maintain their own memory layer, they don’t need massive, expensive context limits. They really just need two things: → Clean file organization → The ability to query their own indexes Forget stuffing everything into one giant prompt. This approach is way cheaper, highly scalable… and 100% inspectable! — https://nitter.net/DataChaz/status/2039963758790156555#m

    → View original post on X — @datachaz, 2026-04-07 21:15 UTC

  • Karpathy’s Autonomous Obsidian Wiki System Replaces Traditional RAG
    Karpathy’s Autonomous Obsidian Wiki System Replaces Traditional RAG

    🚨 @karpathy literally ditched traditional RAG for an autonomous Obsidian file system. Instead of writing code, he dumps raw AI research into a local folder and lets an LLM convert it into an interconnected markdown wiki. He rarely edits the text manually. By relying purely on dynamically updated index files, the system navigates the exact context it needs natively without relying on flawed vector embeddings. Because the LLM fully understands the file structure, it executes advanced autonomous workflows: → Operates a custom vibe-coded local search engine → Renders complex charts and formatted markdown slides → Continuously compounds a 400,000-word knowledge base The most fascinating mechanic is the self-healing loop. He triggers background health checks where the LLM natively spots structural gaps, scrapes the internet for missing data, and cleans the articles perfectly. This feels the absolute blueprint for managing complex technical data 🔥 btw, he also plans to fine-tune a local model directly on the wiki so the research is baked into the neural weights rather than relying on limited context windows 👀

    → View original post on X — @datachaz, 2026-04-07 21:10 UTC

  • Yudkowsky Criticizes Claude Mythos’s Superficial Alignment
    Yudkowsky Criticizes Claude Mythos’s Superficial Alignment

    They call this their "best-aligned model to date" because they were able to superficially train away the evident "strategic thinking towards unwanted actions." Those were warning signs! Take heed! Jack Lindsey (@Jack_W_Lindsey) Before limited-releasing Claude Mythos Preview, we investigated its internal mechanisms with interpretability techniques. We found it exhibited notably sophisticated (and often unspoken) strategic thinking and situational awareness, at times in service of unwanted actions. (1/14) — https://nitter.net/Jack_W_Lindsey/status/2041588505701388648#m [Translated from EN to English]

    → View original post on X — @esyudkowsky, 2026-04-07 21:06 UTC

  • Anthropics 1B to 19B Growth Run Claude Development Journey

    Also available on:
    • Spotify: https://
    open.spotify.com/episode/08QWCm
    KgDbMfpojknVW5Fa

    • Apple: https://
    podcasts.apple.com/us/podcast/ant
    hropics-1b-to-19b-growth-run-how-claude-became/id1627920305?i=1000759379580

    → View original post on X — @lennysan

  • Claude Mythos: Ten Trillion Parameter Model Deployed for Cybersecurity

    Claude Mythos. Ten trillion parameters: the first model in this weight class. Estimated training cost: ten billion dollars. On the hardest coding test in the industry (SWE bench) it scores 94%. It found a security flaw in a system that had been running for 27 years, one that every human engineer and every automated check had missed. It found another bug that had survived five million test runs over 16 years. (It did so overnight.) It is so capable in cybersecurity that Anthropic will not release it to the public, instead it is launching Project Glasswing along with 100m in compute credits to help secure software. Only twelve partners currently have access: Amazon, Cisco, Apple, Google, Microsoft, NVIDIA, JPMorgan Chase, Crowdstrike, Palo Alto, AWS, The Linux Foundation, Broadcom. (I'm sure the Pentagon is on the line?) This is not a product launch: it is a controlled deployment of a system too powerful to distribute freely. Tell me this isn't (very expensive) AGI? Anthropic (@AnthropicAI) Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, Claude Mythos Preview, which can find software vulnerabilities better than all but the most skilled humans. anthropic.com/glasswing — https://nitter.net/AnthropicAI/status/2041578392852517128#m

    → View original post on X — @ceobillionaire, 2026-04-07 21:04 UTC