We have a position paper led by the awesome @_shruti_joshi_ and @rpatrik96 that shows how causality can provide a unifying framework to formalize, estimate and evaluate interpretability methods for foundation models. Have a look! https://t.co/pcnI4XzHUk
— Dhanya Sridhar (@dhanya_sridhar) 24 mars 2026
We have a position paper led by the awesome @_shruti_joshi_ and @rpatrik96 that shows how causality can provide a unifying framework to formalize, estimate and evaluate interpretability methods for foundation models. Have a look! Shruti Joshi (@_shruti_joshi_) Mechanistic interpretability aims to understand models — and the more superhuman or incoherent they become, the more we need that understanding to be reliable. We propose a framework for this, drawing on established tools from causal reasoning and statistical identifiability: 🧵 — https://nitter.net/_shruti_joshi_/status/2035025756632302039#m
→ View original post on X — @hugo_larochelle, 2026-03-24 16:12 UTC