Abstract
Key-Value cache (\texttt\{KV\} \texttt\{cache\}) compression has emerged as a promising technique to optimize Large Language Model (LLM) serving. It primarily decreases the memory consumption of \texttt\{KV\} \texttt\{cache\} to reduce the computation cost. Despite the development of many compression algorithms, their applications in production environments are still not prevalent. In this paper,