"When implementations of the Transformer's self-attention layer utilize SRAM instead of DRAM, they can achieve significant speedups." Thanks @moritzthuening
, read the full paper here –>
SRAM Optimization Accelerates Transformer Self-Attention Performance
By
–