Instead of deleting tokens or optimizing a new prompt-compressed cache, Still learns to synthesize a compact KV cache in a single forward pass. Thus, a small Perceiver per layer reads the cache.
Still: Amortized KV Cache Compaction in a Single Forward Pass
By
–
