A really cool paper from @Zai_org "IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse" This paper shows that in sparse-attention LLMs, nearby layers usually pick almost the same important tokens. So you can cache and reuse those token indices instead of
IndexCache Accelerates Sparse Attention via Cross-Layer Index Reuse
By
–
