← all papers · overview

Layer-wise Quantization: A Pragmatic And Effective Method For Quantizing Llms Beyond Integer Bit-levels

Abstract

We present a simple meta quantization approach that quantizes different layers of a large language model (LLM) at different bit levels, and is independent of the underlying quantization technique. Specifically, we quantize the most important layers to higher bit precision and less important layers to lower bits. We propose two effective strategies to measure the importance of layers within LLMs: t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).