The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit paper page: https://
huggingface.co/papers/2306.17
759
… In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network's trainability. Motivated by the success of
Shaped Transformer: Attention Models in Infinite Depth-Width Limit
By
–
