DOI: 10.1145/3837104 ISSN: 2836-6573

CatKV: Accelerating Position-Independent Caching with Context-Adaptive KV Cache Compression

Jie Xu, Renjie Liu, Yi Li, Bo Tang

Context caching is widely adopted to reduce the dominant prefill cost for long prompts across various large language model (LLM) applications. Recently, position-independent caching (PIC) has emerged to improve KV cache reuse beyond strict prefix matching. However, the required KV caches often reside in a persistent remote store in realistic deployments, making KV transfer the new bottleneck on time-to-first-token (TTFT). Existing KV compression techniques are primarily designed for long contexts. Since text chunks in PIC are much smaller, they fail to deliver sufficient compression and/or lead to substantial generation quality degradation.

In this work, we propose CatKV, an end-to-end PIC serving system for efficient KV cache reuse. Specifically, we devise a context-adaptive mixed-precision compression scheme in CatKV, which combines SVD factorization with adaptive scalar quantization in latent space. Given a target compression ratio, CatKV automatically determines the SVD and quantization configurations by analyzing the singular-value spectrum using the elbow method. To improve system efficiency, CatKV offloads online compression of cold text chunks to the CPU for asynchronous execution and exploits grouped factorization across chunks to further reduce storage consumption. Furthermore, CatKV introduces a six-stage layer-wise pipeline to reduce end-to-end latency. Extensive experiments demonstrate the superiority of CatKV across 3 PIC algorithms, 4 benchmarks, and 3 LLMs. In particular, CatKV achieves 2–5× faster TTFT and 6–10× higher compression ratios compared to baselines, all while preserving model accuracy. CatKV is open-source at https://github.com/AlayaDB-AI/CatKV.git.