← all papers · overview

Rethinking Key-value Cache Compression Techniques For Large Language Model Serving

Abstract

Key-Value cache (\texttt\{KV\} \texttt\{cache\}) compression has emerged as a promising technique to optimize Large Language Model (LLM) serving. It primarily decreases the memory consumption of \texttt\{KV\} \texttt\{cache\} to reduce the computation cost. Despite the development of many compression algorithms, their applications in production environments are still not prevalent. In this paper,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).