← all papers · overview

Turboboa: Faster And Exact Attention-aware Quantization Without Backpropagation

Abstract

The rapid growth of large language models (LLMs) has heightened the importance of post-training quantization (PTQ) for reducing memory and computation costs. Among PTQ methods, GPTQ has gained significant attention for its efficiency, enabling billion-scale LLMs to be quantized within a few GPU hours. However, GPTQ's assumption of layer-wise independence leads to severe accuracy drops in low-bit r

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).