← all papers · overview

CSKV: Training-efficient Channel Shrinking For KV Cache In Long-context Scenarios

Abstract

Large Language Models (LLMs) have been widely adopted to process long-context tasks. However, the large memory overhead of the key-value (KV) cache poses significant challenges in long-context scenarios. Existing training-free KV cache compression methods typically focus on quantization and token pruning, which have compression limits, and excessive sparsity can lead to severe performance degradat

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).