AI Dynamics

Global AI News Aggregator

About

Massive attention matrices cause Transformer interpretability crisis

Here's the interpretability crisis no one talks about: Every Transformer layer creates an L×L attention matrix. Across 96 layers and 96 heads in GPT-4, that's a MASSIVE evolving tensor cloud. The paper proves this is why we can't interpret these models not their size, but their

→ View original post on X — @godofprompt