LRU is harder to beat than the KV-cache papers suggest

Real-world agentic LLM serving traces show that the default LRU policy for KV/prefix caches is much harder to beat than recent papers suggest, even under capacity pressure. For CIOs and technology leaders, the strategic takeaway is that service quality and cost efficiency will likely be improved more by right-sizing cache capacity, understanding workload trace patterns, and reducing recomputation in tight tool-call loops than by betting on exotic eviction algorithms or TTL-based assumptions. For IT organizations, this means cache optimization should be driven by production telemetry and simulator validation against real traces, not by paper benchmarks alone.

Hacker News3 min read
Read full article
LRU is harder to beat than the KV-cache papers suggest

Read the full story at Hacker News →