← all papers · overview

More Than A Quick Glance: Overcoming The Greedy Bias In Kv-cache Compression

Abstract

While Large Language Models (LLMs) can theoretically support extensive context windows, their actual deployment is constrained by the linear growth of Key-Value (KV) cache memory. Prevailing compression strategies mitigate this through various pruning mechanisms, yet trade-off semantic recall for memory efficiency. In this work, we present LASER-KV (Layer Accumulated Selection with Exact-LSH Recal

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).