← all papers · overview

Dquant: Accurate Low-bit Post-training Weight Quantization For Llms

Abstract

Large language models (LLMs) deliver strong performance, but their high compute and memory costs make deployment difficult in resource-constrained scenarios. Weight-only post-training quantization (PTQ) is appealing, as it reduces memory usage and enables practical speedup without low-bit operators or specialized hardware. However, accuracy often degrades significantly in weight-only PTQ at sub-4-

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).