← all papers · overview

Bisup: Bidirectional Quantization Error Suppression For Large Language Models

Abstract

As the size and context length of Large Language Models (LLMs) grow, weight-activation quantization has emerged as a crucial technique for efficient deployment of LLMs. Compared to weight-only quantization, weight-activation quantization presents greater challenges due to the presence of outliers in activations. Existing methods have made significant progress by exploring mixed-precision quantizat

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).