AI Dynamics

Global AI News Aggregator

About

Prefill-as-Service Enables Cross-Datacenter LLM Serving

new paper from Kimi! "Prefill-as-a-Service makes long-context LLM serving cross-datacenter" Main idea: smaller KV Cache turns prefill into a scalable remote service. Prefill is compute-bound, decode is memory-bandwidth-bound, but KV cache transfer keeps them trapped in the

→ View original post on X — @askalphaxiv