8/ Patchscopes – proposes a framework that leverages a model itself to explain its internal representations; it can be used to answer questions about an LLM’s computation and can even be used to fix latent multi-hop reasoning errors.
Patchscopes: Framework for Explaining LLM Internal Representations
By
–
