"GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent" Instead of giving LLMs huge KV caches or one-shot compressed summaries, GradMem shows that you can let a frozen model take some test-time gradient steps to write a long context into a small memory.
GradMem: Gradient Descent Memory for Context Compression
By
–
