If AGI is achievable & labs can be banned from using a model internally ONLY if they release the model publicly, the Big Three labs may decide it is better to capture all the value from AGI themselves by expansion & acquisition. Sharing AI access with other firms triggers risk.
SAFETY
-
AI finding security bugs threatens code writing use case
By
–
We won't know until/unless we hear what the bug actually was If it really was "asking Fable to find security bugs in code found security bugs" then it can't be fixed without making the model useless for writing any code at all, which is one of the most important use-cases
-

Why Collaborating with Regulators Ensures AI Ethics
By
–
Why Collaborating with Regulators Ensures #AIEthics by @antgrasso #ArtificialIntelligence #MachineLearning #ML
-
OrcaRouter detects hallucinations and ensures reliability
By
–
A model has a bad day and your output collapses.
OrcaRouter makes reliability structural, not a matter of luck.
If a leg hallucinates, the judge detects it and provides the answer that holds the road. -
Anthropic’s double standard on AI security and export controls
By
–
Anthropic: Red alert! Our models are a major security risk!
Also Anthropic: How could the government slap export controls on them? Outrageous! -

Problem Detection in Agent Traces with Post-Trained Model
By
–
Detecting problems in agent traces in production is difficult. It must be done at low cost (due to volume) but also with accuracy (otherwise too much noise). We post-trained our own model for this. SOTA accuracy, at ~10-100x lower cost.
-
AI may never be jailbreak-proof nor hallucination-free
By
–
And AI systems may never be jailbreak-proof or hallucination free. And individual queries may matter less than a bad actor breaking a problem down into pieces and feeding it through multiple projects and prompts. And AIs themselves may change behavior unpredictably with context.
-
Complexity of AI Regulation: Models Are Just One Piece
By
–
Bright regulatory lines for AI are inherently complicated because models are just a piece of the puzzle: harnesses can make models more capable, a less capable open system may be more or less riskier than a more capable closed one, skills/connected systems change risk levels, etc
-

New video: analysis of political conflict and Fable 5 blockade
By
–
NEW VIDEO in the LAB! Today, analyzing the political conflict that led to the first government blockade of a frontier model, in this case, Fable 5. Reasons, consequences, and my opinion. All in the video. Link below
-
Mandatory verifiable KYC implementation for Claude API users
By
–
So now every company that builds on top of the Claude API also needs to implement KYC, in a way that can then be communicated to Anthropic such that they can trust it?