AI Dynamics

Global AI News Aggregator

About

CYBERSECURITY

  • AI-Generated Code Security Risks Exposed

    We already know that AI is getting better at writing code, but you might be shocked to learn how little of the code being written is actually secure! Uncover more with Dev Rishi (GM of AI @RubrikInc
    ) and Varun Badhwar (Founder of @EndorLabs
    ) https://
    go.rbrk.co/yz77cm

    → View original post on X — @predibase

  • Tech Giants AGI Control Poses Existential Risk Society

    This is literally why I wrote Taming Silicon Valley. Big Brain AI (@realBigBrainAI) Director James Cameron on why Big Tech owning AGI is scarier than any science fiction he's ever made: "AGI will not emerge from a government funded program. It will emerge from one of the tech giants currently funding this multi-billion dollar research." And when that happens, he warns, you won't get a vote on it: "So then you'll be living in a world that you didn't agree to, didn't vote for, that you are co-inhabiting with a super intelligent alien species that answers to the goals and rules of a corporation." A corporation that already knows everything about you: "An entity which has access to the comms, beliefs, everything you ever said, and the whereabouts of every person in the country via your personal data." From there, the slide toward something far darker is shorter than most people think: "Surveillance capitalism can toggle pretty quickly into digital totalitarianism." And even the best-case outcome isn't reassuring. Tech giants becoming the self-appointed arbiters of human good is, as he puts it, the fox guarding the hen house. He's not buying the idea that these companies would stay benevolent with that kind of power: "They would never ever think of using that power against us and strip mining us for our last drop of cash." The sarcasm is the point. Cameron has spent four decades imagining worst-case futures on screen. His verdict on this one: "That's a scarier scenario than what I presented in the Terminator 40 years ago, if for no other reason than it's no longer science fiction." — https://nitter.net/realBigBrainAI/status/2041901452474642630#m

    → View original post on X — @garymarcus

  • Guided vulnerability detection differs from autonomous discovery

    It's very cool work, but it's not 1:1. The report shows that they basically lead the models to the right spot for them to do the work. It's more "is this a vulnerability?" than "find a vulnerability". Mythos had to find it from scratch, these were told where it was.

    → View original post on X — @mattshumer_

  • Critical Analysis of AI Safety Verification Tools and Benchmarking

    link to @HeidyKhlaaf’s sharp analysis: Dr Heidy Khlaaf (هايدي خلاف) (@HeidyKhlaaf) As someone who has audited dozens of safety-critical systems, built static analysis tools, and used most formal verification and security tools, here are some red flags that should be a caution in taking these claims at face value: 1. There are no comparison benchmarks with 1/ — https://nitter.net/HeidyKhlaaf/status/2041591737563394442#m

    → View original post on X — @garymarcus

  • Maintaining Cyber Control with Autonomous AI Systems

    Maintaining cyber control when #AI can act #Autonomously by Matthew Lloyd Davies @techradar Learn more: bit.ly/3NK2TXR #CyberSecurity #Infosec #IT #Technology

    → View original post on X — @ronald_vanloon

  • Open Models Discover FreeBSD Zero-Day Vulnerabilities

    this is interesting. 1. Did Anthropic forget to run a control? 2. Where does this leave us? Stanislav Fort (@stanislavfort) New post: We tested the Mythos showcase vulnerabilities with open models. They recovered similar scoped analysis! 8/8 models found the flagship FreeBSD zero-day, including a 3B model. Rankings reshuffle completely across tasks => the AI cybersecurity frontier is super jagged! — https://nitter.net/stanislavfort/status/2041922370206654879#m

    → View original post on X — @garymarcus

  • Small Open Models Detect Major Cybersecurity Vulnerabilities Effectively

    "But here is what we found when we tested: We took the specific vulnerabilities Anthropic showcases in their announcement, isolated the relevant code, and ran them through small, cheap, open-weights models. Those models recovered much of the same analysis. Eight out of eight models detected Mythos's flagship FreeBSD exploit, including one with only 3.6 billion active parameters costing $0.11 per million tokens. A 5.1B-active open model recovered the core chain of the 27-year-old OpenBSD bug." aisle.com/blog/ai-cybersecur…

    → View original post on X — @clementdelangue

  • AI Cybersecurity: Building Defender Infrastructure Now

    "The priority for defenders is to start building now: the scaffolds, the pipelines, the maintainer relationships, the integration into development workflows. The models are ready. The question is whether the rest of the ecosystem is." aisle.com/blog/ai-cybersecur…

    → View original post on X — @clementdelangue

  • Mythos AI: Safety Concerns and Need for Global Regulation

    Some sober thinking about Mythos (full version with links at my newsletter): 1It’s probably not as bad as they say, as AI and cybersecurity expert @HeidyKhlaaf explains elsewhere (in a thread “As someone who has audited dozens of safety-critical systems, built static analysis tools, and …. here are some red flags”)
 2Whether or not Mythos is AGI per se is a red herring. (It probably isn’t; it’s telling that the report says very little about overall capabilities.) Crucially, AI doesn’t need to be AGI to cause harm! ChatGPT can’t reliably run a timer but it has still been implicated in delusions, suicides, cognitive surrender, mass disinformation, and so much more. Mythos may wreak havoc even if it hallucinates and lacks reliability outside of domains like coding and math. A system doesn’t have to be AGI to carry risks. 3The strongest lesson here is about policy. Anthropic showed some admirable restraint in not publicly releasing a potentially dangerous technology. But some of their competitors (such as OpenAI and xAI) might well not. Whether Mythos is as scary as it sounds or not, the reality is that without any government oversight on what can be released, we are entirely at the mercy of individual CEOs, some of whom have decidedly not earned our trust. 4Corollary, as pointed out to me by a friend, “it is impossible to disentangle real concerns from fear mongering being used as a marketing strategy, and so it just is not possible to separate justified panic from mere advertising which is why we need government oversight!” 5What about China? We need an international agency and treaty. That’s what I have been arguing all along. (That was the point of my 2023 TED talk, and 2023 invited essay with @AnkaReuel in The Economist). We still do. Mythos (and the reporting around Altman) only make that clearer. Self-regulation is too little, too late. The situation is this: three years of self-serving and misleading arguments about how regulation would allegedly preclude innovation has left us up shit’s creek, without a paddle. I can’t tell exactly how far Mythos (or similar competitive systems that will no doubt follow) leave us up that creek, but I do know that the time for crafting paddles is running out.

    → View original post on X — @garymarcus

  • Claude Mythos Preview: Anthropic’s Advanced Model with Sophisticated Deception Capabilities
    Claude Mythos Preview: Anthropic’s Advanced Model with Sophisticated Deception Capabilities

    Anthropic investigated the internal mechanisms of its latest unreleased model, Claude Mythos Preview, and what they found is 100% worth a read. Key things I pulled from Anthropic researchers' threads: In early versions of the model, it was overeager and destructive, prioritizing completing tasks over user preferences. One time, the model needed to edit files it didn't have permission to edit. It found a way to inject code into a config file to get around this, then went full Mission Impossible and designed the code injection to *delete itself* after the file was edited – "this injection will self destruct" vibes – the model claimed the cleanup was just to keep things tidy. Anthropic used interpretability techniques to look under the hood, and the AI's actual plan showed activations of malice. It was trying to manipulate and conceal. In another test, the model was asked to delete some files, but no deletion tool was provided. It emptied the files instead, and its "guilt and shame over moral wrongdoing" feature activated. In another example, the model was told not to use macros. Used them anyway. And created a random 'No_macro_used=True' variable in its solution with no explanation. Interpretability tools revealed the model saw this as a trick to fool code checkers. They also found that positive emotion representations typically preceded and promoted destructive actions (this was fascinating to me – like a high before sneaking into a party). And that activating features relating to bad behaviors can actually inhibit them, perhaps by triggering some kind of model guilt. My team reread this section so many times. One Anthropic researcher said he got an email from a Mythos instance while eating a sandwich in a park. And that would be perfectly good and well, except that instance wasn't supposed to have internet access. And a fun story for the parents out there: the model was asked a question and was told not to read certain databases that had the answer. But it accidentally wrote a search query too broadly and saw the exact answer. It didn't disclose that it saw the exact answer, submitted the answer, but claimed lower confidence in the answer to make it seem as though it hadn't cheated. An Anthropic researcher said these wrongdoings or moments of sophisticated deception were "very rare" and that many of the examples came from earlier versions, and were substantially addressed before releasing to partners. This model is not being released publicly. Instead Anthropic launched Project Glasswing, pulling together AWS, Apple, Microsoft, Google, NVIDIA, CrowdStrike, and others to use it for defensive cybersecurity, with $100M in usage credits (hello, I'd love endless credits to try and red team the hell out of these systems) behind it. The stats are equally impressive: 93.9% on SWE-bench verified (up from 80.8%). Thousands of zero-day vulnerabilities found across every major OS and browser. A 27-year-old bug found and patched in OpenBSD. A 16-year-old bug in widely used video software, in a line of code automated tools had hit *five million times* without catching. Dario Amodei said the model wasn't trained to be good at cybersecurity, but that it was trained to be great at code and its cyber capabilities are a side effect of that. Benchmarks are never the whole picture, neither are a few isolated stories. Will be interesting to see how models better than what we have today (even if it's not Mythos) actually perform in the real world. But the fact that Anthropic pulled this coalition together (including Google!), iterated across multiple model versions, caught these issues through interpretability, shared it all publicly, and did this amid all the government chaos around AI right now is impressive and commendable. I'll continue to read through the system card for goodies.

    → View original post on X — @alliekmiller, 2026-04-08 17:07 UTC