← all papers · overview

Uncomp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design From An Uncertainty-aware Perspective

Abstract

Deploying large language models (LLMs) for long-context inference remains challenging due to their substantial memory and computational demands. While techniques such as Key-Value (KV) cache compression are designed to reduce memory usage, they often neglect the structured sparsity inherent in the relationship between hidden states and their corresponding KV cache. In this work, we explore the rol

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).