What if you could reduce the memory of a language model by 50x in seconds without losing performance? Researchers from MIT present Fast KV Compaction via Attention Matching. They build compact key-value caches in the latent space that preserve
Fast KV Compaction reduces LLM memory by 50x without loss
By
–
