DOI: 10.1145/3837101 ISSN: 2836-6573

Buoy: Efficient and Effective Cache Replacement for Prefix Caching

Liang Wang, Ranjun Jia, Kai Wang, Kai Lu, Jiguang Wan

Modern large language model (LLM) systems widely employ prefix caching to enable key-value (KV) cache reuse across different queries to minimize inference costs. At the heart of prefix caching is the replacement algorithm, which is crucial for managing limited cache space across the GPU–CPU memory hierarchy. However, the prevalent least-recently-used (LRU) algorithm is not prefix-aware or tier-aware, and remains far from optimal. In this paper, we present Buoy, a novel and general cache replacement policy for prefix caching. Buoy achieves quick demotion and considers prefix-matched access characteristics to improve cache efficiency (hit rate). Meanwhile, Buoy enables effective cache tiering to reduce access latency upon cache hits. We evaluate Buoy on real-world traces, and the experimental results demonstrate its superiority in improving hit rates and prefill throughput over state-of-the-art algorithms. Compared to LRU, Buoy achieves 6–93% higher cache hit rates and increases LLM prefill throughput by 8–16%.