Don't trust me – look at the data!
SAFETY
-
Vibe Coding Definition Debate: Unreviewed LLM-Generated Code
By
–
I will stubbornly continue to use the original definition! https://
simonwillison.net/2025/Mar/19/vi
be-coding/
… Using it to mean any form of AI-assisted programming deeply frustrates me because we really need a term that means "LLM-written code that nobody ever took the time to review" -
Anthropic Hiring Research Engineer Alignment Science Team
By
–
If you’re interested in joining us to work on these and related issues, you can apply for our Research Engineer/Scientist role (
https://
job-boards.greenhouse.io/anthropic/jobs
/4631822008
…) on the Alignment Science team. -
Classifiers Detect Misalignment Risks and CBRN Threats
By
–
There’s plenty of work to be done to make the classifiers even more accurate and effective. In the future, they might even be able to remove data relevant to misalignment risks (scheming, deception, and so on), as well as CBRN risks.
-
Claude 3 Sonnet Classifier Detects CBRN Information in Training Data
By
–
We trained six different classifiers to detect and remove CBRN information from training data. The best and most efficient results were from a classifier that used a small model from the Claude 3 Sonnet series to flag the harmful data.
-

CBRN Filtering Reduces Harmful Capabilities Without Affecting Science
By
–
One concern is that filtering CBRN data will reduce performance on other, harmless capabilities—especially science. But we found a setup where the classifier reduced CBRN accuracy by 33% beyond a random baseline with no particular effect on a range of other benign tasks.
-
AI Training Data Filtering Removes Hazardous Information
By
–
The wealth of data used in AI training contains hazardous CBRN information. Developers usually train models not to use it. Here, we tried removing the information at the source, so even if models are jailbroken, the info isn't available. Read more: https://
alignment.anthropic.com/2025/pretraini
ng-data-filtering/
… -
Darwin Gödel Machine: Self-Improving AI System Breakthrough
By
–
Science fiction? Far from it. Researchers recently created the Darwin Gödel Machine, which is “a self-improving system that iteratively modifies its own code.” https://
bit.ly/4mvLKND #AI #selfevolvingAI #ArtificialIntelligence #Technology #TechNews -
Small models could topple AI scaling narrative
By
–
If small language models outperform LLMs in practice, the whole AI scaling narrative might collapse overnight.