9/ ThinK – proposes an approach to address inefficiencies in KV cache memory consumption; it focuses on the long-context scenarios and the inference side of things; it presents a query-dependent KV cache pruning method to minimize attention weight loss while selectively pruning
ThinK: Query-Dependent KV Cache Pruning for Efficiency
By
–