AI Dynamics

Global AI News Aggregator

About

DMC Improves Transformer Inference Speed Through Context Compression

when contexts are long, attending to every single token in the past feels wasteful (and not at all how human brains work). feels like a natural setting for compression… DMC seems like a huge improvement in transformer inference speed — congrats to the authors!

→ View original post on X — @jxmnop