AI Dynamics

Global AI News Aggregator

About

Anthropic’s Model Diffing Reveals Post-Training Effects on LLMs

just learned about "model diffing" from Anthropic. buried in an october blogpost; feels really novel. training a 'crosscoder' between two models of the same family produces interpretable diffs. here post-training clearly adds refusals, QA, math, etc. pretty amazing stuff

→ View original post on X — @jxmnop