← all papers · overview

Wkvquant: Quantizing Weight And Key/value Cache For Large Language Models Gains More

Abstract

Large Language Models (LLMs) face significant deployment challenges due to their substantial memory requirements and the computational demands of auto-regressive text generation process. This paper addresses these challenges by focusing on the quantization of LLMs, a technique that reduces memory consumption by converting model parameters and activations into low-bit integers. We critically analyz

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).