← all papers · overview

Prefixing Attention Sinks Can Mitigate Activation Outliers For Large Language Model Quantization

Abstract

Despite recent advances in LLM quantization, activation quantization remains to be challenging due to the activation outliers. Conventional remedies, e.g., mixing precisions for different channels, introduce extra overhead and reduce the speedup. In this work, we develop a simple yet effective strategy to facilitate per-tensor activation quantization by preventing the generation of problematic tok

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).