← all papers · overview

VQKV: High-fidelity And High-ratio Cache Compression Via Vector-quantization

Abstract

The growing context length of Large Language Models (LLMs) enlarges the Key-Value (KV) cache, limiting deployment in resource-limited environments. Prior training-free approaches for KV cache compression typically rely on low-rank approximation or scalar quantization, which fail to simultaneously achieve high compression ratios and high reconstruction fidelity. We propose VQKV, a novel, training-f

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).