← all papers · overview

QLLM: Accurate And Efficient Low-bitwidth Quantization For Large Language Models

Abstract

Large Language Models (LLMs) excel in NLP, but their demands hinder their widespread deployment. While Quantization-Aware Training (QAT) offers a solution, its extensive training costs make Post-Training Quantization (PTQ) a more practical approach for LLMs. In existing studies, activation outliers in particular channels are identified as the bottleneck to PTQ accuracy. They propose to transform t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).