← all papers · overview

Finequant: Unlocking Efficiency With Fine-grained Weight-only Quantization For Llms

Abstract

Large Language Models (LLMs) have achieved state-of-the-art performance across various language tasks but pose challenges for practical deployment due to their substantial memory requirements. Furthermore, the latest generative models suffer from high inference costs caused by the memory bandwidth bottleneck in the auto-regressive decoding process. To address these issues, we propose an efficient

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).