← all papers · overview

No Token Left Behind: Reliable KV Cache Compression Via Importance-aware Mixed Precision Quantization

Abstract

Key-Value (KV) Caching has become an essential technique for accelerating the inference speed and throughput of generative Large Language Models~(LLMs). However, the memory footprint of the KV cache poses a critical bottleneck in LLM deployment as the cache size grows with batch size and sequence length, often surpassing even the size of the model itself. Although recent methods were proposed to s

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).