← all papers · overview

Kvsharer: Efficient Inference Via Layer-wise Dissimilar KV Cache Sharing

Abstract

The development of large language models (LLMs) has significantly expanded model sizes, resulting in substantial GPU memory requirements during inference. The key and value storage of the attention map in the KV (key-value) cache accounts for more than 80% of this memory consumption. Nowadays, most existing KV cache compression methods focus on intra-layer compression within a single Transformer l

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).