A new paper just exposed how much money AI labs waste on GPUs. Running LLMs at scale hits a wall when prefill and decode share one datacenter. The KV cache transfer between them is massive. This forces expensive RDMA networks and identical hardware everywhere. Kimi proposes
Kimi Proposes Solution to LLM Prefill-Decode KV Cache Transfer Problem
By
–
