Prefill-as-a-Service:KVCache of Next-Generation Models Could Go Cross-Datacenter
Prefill-as-a-Service (PrfaaS) enables large language model inference to be distributed across geographically separated datacenters by selectively offloading prefill processing to specialized clusters and transferring compressed KVCache over standard networks, achieving 54% higher throughput than traditional single-cluster architectures. This breakthrough decouples prefill and decode infrastructure, allowing IT organizations to independently scale compute resources across multiple datacenters while reducing reliance on expensive, low-latency RDMA fabrics. The strategic implication is significant cost reduction and operational flexibility for enterprises deploying large-scale AI workloads, enabling heterogeneous hardware utilization and dynamic resource allocation across loosely coupled infrastructure.