← all papers · overview

A Comprehensive Evaluation Of Quantization Strategies For Large Language Models

Abstract

Increasing the number of parameters in large language models (LLMs) usually improves performance in downstream tasks but raises compute and memory costs, making deployment difficult in resource-limited settings. Quantization techniques, which reduce the bits needed for model weights or activations with minimal performance loss, have become popular due to the rise of LLMs. However, most quantizatio

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).