← all papers · overview

Keepkv: Achieving Periodic Lossless KV Cache Compression For Efficient LLM Inference

Abstract

Efficient inference of large language models (LLMs) is hindered by an ever-growing key-value (KV) cache, making KV cache compression a critical research direction. Traditional methods selectively evict less important KV cache entries, which leads to information loss and hallucinations. Recently, merging-based strategies have been explored to retain more information by merging KV pairs that would b

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).