← all papers · overview

Identify Critical KV Cache In LLM Inference From An Output Perturbation Perspective

Abstract

Large language models have revolutionized natural language processing but face significant challenges of high storage and runtime costs, due to the transformer architecture's reliance on self-attention, particularly the large Key-Value (KV) cache for long-sequence inference. Recent efforts to reduce KV cache size by pruning less critical entries based on attention weights remain empirical and lack

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).