when contexts are long, attending to every single token in the past feels wasteful (and not at all how human brains work). feels like a natural setting for compression… DMC seems like a huge improvement in transformer inference speed — congrats to the authors!
DMC Improves Transformer Inference Speed Through Context Compression
By
–
