new paper from Kimi! "Prefill-as-a-Service makes long-context LLM serving cross-datacenter" Main idea: smaller KV Cache turns prefill into a scalable remote service. Prefill is compute-bound, decode is memory-bandwidth-bound, but KV cache transfer keeps them trapped in the
Prefill-as-Service Enables Cross-Datacenter LLM Serving
By
–
