← all papers · overview

Sliderquant: Accurate Post-training Quantization For Llms

Abstract

In this paper, we address post-training quantization (PTQ) for large language models (LLMs) from an overlooked perspective: given a pre-trained high-precision LLM, the predominant sequential quantization framework treats different layers equally, but this may be not optimal in challenging bit-width settings. We empirically study the quantization impact of different layers on model accuracy, and ob

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).