← all papers · overview

Outliertune: Efficient Channel-wise Quantization For Large Language Models

Abstract

Quantizing the activations of large language models (LLMs) has been a significant challenge due to the presence of structured outliers. Most existing methods focus on the per-token or per-tensor quantization of activations, making it difficult to achieve both accuracy and hardware efficiency. To address this problem, we propose OutlierTune, an efficient per-channel post-training quantization (PTQ)

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).