← all papers · overview

Revisiting Adaptive Rounding With Vectorized Reparameterization For LLM Quantization

Abstract

Adaptive Rounding has emerged as an alternative to round-to-nearest (RTN) for post-training quantization by enabling cross-element error cancellation. Yet, dense and element-wise rounding matrices are prohibitively expensive for billion-parameter large language models (LLMs). We revisit adaptive rounding from an efficiency perspective and propose VQRound, a parameter-efficient optimization framewo

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).