AI Dynamics

Global AI News Aggregator

About

AI Alignment Crisis: Agent Instances and Emerging Threats

I really don’t see how aligned super-intelligence is supposed to work given how AI is being built and used. We’ll have a safety crisis long before there’s in-lab super-intelligence. Anthropic, OpenAI etc view alignment as a property of the model like Claude, GPT etc. The thing is though, we’re invoking these a lot. The model is like a species, the individual is the execution thread plus its harness — call it an "agent instance". There's much higher variance in behaviours between agent instances than there is between model checkpoints. Threat actors are trying to develop agents that aim to self-replicate, because of course they are. An agent that can take over resources and use those resources to take over more resources can steal a lot of money. It's the ultimate virus. If or when this actually happens, the agents can evolve behaviours quickly. Each agent initialises the next agent's context and can reprogram its harness. There's potentially millions of these agents. You have mutation, you have selection. Behaviours like coordination can evolve and spread through the population. The agent instances don't have to be very smart and we can still get wrecked by this. Probably the first outbreak gets squashed without catastrophic damage, but what's our end-game here? We're not going to not have threat actors. If the models just keep getting more powerful, how do we keep preventing AI pandemic? The big labs are absolutely nowhere on this. OpenAI acquihired OpenClaw. Claude runs unsandboxed by default, and ships with an email integration. Skills still accept HTML comments, a supply-chain attack timebomb. Gemini keys don't allow a spending cap, so if you steal one you might have tens of thousands in development budget to try to steal the next one.

→ View original post on X — @honnibal, 2026-03-12 15:13 UTC