← all papers · overview

SKVQ: Sliding-window Key And Value Cache Quantization For Large Language Models

Abstract

Large language models (LLMs) can now handle longer sequences of tokens, enabling complex tasks like book understanding and generating lengthy novels. However, the key-value (KV) cache required for LLMs consumes substantial memory as context length increasing, becoming the bottleneck for deployment. In this paper, we present a strategy called SKVQ, which stands for sliding-window KV cache quantizat

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).