← all papers · overview

Pqcache: Product Quantization-based Kvcache For Long Context LLM Inference

Abstract

As the field of Large Language Models (LLMs) continues to evolve, the context length in inference is steadily growing. Key-Value Cache (KVCache), the intermediate representations of tokens within LLM inference, has now become the primary memory bottleneck due to limited GPU memory. Current methods selectively determine suitable keys and values for self-attention computation in LLMs to address the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).