This paper just murdered the foundation of every AI model you've ever used. A researcher proved you can match Transformer performance WITHOUT computing a single attention weight. Here's what changed (and why this matters now):
Transformer performance matched without computing attention weights
By
–
