The "Attention is all you need paper" that introduced Transformer neural network architecture has been around for roughly six years. It's by far the first architecture to maintain its universality for a long time, not just for a single modality but for other modalities as well.
Transformers: Six Years of Universal Neural Network Architecture
By
–
