← all papers · overview

CLAQ: Pushing The Limits Of Low-bit Post-training Quantization For Llms

Abstract

Parameter quantization for Large Language Models (LLMs) has attracted increasing attentions recently in reducing memory costs and improving computational efficiency. Early approaches have been widely adopted. However, the existing methods suffer from poor performance in low-bit (such as 2 to 3 bits) scenarios. In this paper, we present a novel and effective Column-Level Adaptive weight Quantizatio

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).