AI Dynamics

Global AI News Aggregator

About

Shaped Transformer: Attention Models in Infinite Depth-Width Limit

The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit paper page: https://
huggingface.co/papers/2306.17
759
… In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network's trainability. Motivated by the success of

→ View original post on X — @_akhaliq