1) Sparse Attention It limits the attention computation to a subset of tokens by: – Using local attention (tokens attend only to their neighbors).
– Letting the model learn which tokens to focus on. But this has a trade-off between computational complexity and performance.
Sparse Attention: Local, Learned Focus with Trade-off
By
–
