← all papers · overview

Model Tells You Where To Merge: Adaptive KV Cache Merging For Llms On Long-context Tasks

Abstract

How to efficiently serve Large Language Models (LLMs) has become a pressing issue because of their huge computational cost in their autoregressive generation process. To mitigate computational costs, LLMs often employ the KV Cache technique to improve the generation speed. While improving the computational efficiency, the storage requirements of the KV cache are substantial, particularly in long-c

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).