← all papers · overview

What Makes Quantization For Large Language Models Hard? An Empirical Study From The Lens Of Perturbation

Abstract

Quantization has emerged as a promising technique for improving the memory and computational efficiency of large language models (LLMs). Though the trade-off between performance and efficiency is well-known, there is still much to be learned about the relationship between quantization and LLM performance. To shed light on this relationship, we propose a new perspective on quantization, viewing it

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).