"Exclusive Self Attention" This paper proposed Exclusive Self-Attention (XSA), which is a tiny two-line change that stops attention from looking at itself. This forces it to focus on the rest of the sequence, and can make transformers more effective! This improves the
Exclusive Self-Attention: Two-Line Transformer Improvement
By
–
