The line that stayed with me: semantic structure may be useful for function without being driven by function. That is a powerful caution for mechanistic interpretability. Some beautiful structures inside models may be less like “designed concepts” and more like the linear
Semantic structure vs. function: A caution for mechanistic interpretability
By
–