← all papers · overview

Regularized Calibration With Successive Rounding For Post-training Quantization

Abstract

Large language models (LLMs) deliver robust performance across diverse applications, yet their deployment often faces challenges due to the memory and latency costs of storing and accessing billions of parameters. Post-training quantization (PTQ) enables efficient inference by mapping pretrained weights to low-bit formats without retraining, but its effectiveness depends critically on both the qua

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).