← all papers · overview

Smoothquant+: Accurate And Efficient 4-bit Post-training Weightquantization For LLM

Abstract

Large language models (LLMs) have shown remarkable capabilities in various tasks. However their huge model size and the consequent demand for computational and memory resources also pose challenges to model deployment. Currently, 4-bit post-training quantization (PTQ) has achieved some success in LLMs, reducing the memory footprint by approximately 75% compared to FP16 models, albeit with some acc

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).