← all papers · overview

AFPQ: Asymmetric Floating Point Quantization For Llms

Abstract

Large language models (LLMs) show great performance in various tasks, but face deployment challenges from limited memory capacity and bandwidth. Low-bit weight quantization can save memory and accelerate inference. Although floating-point (FP) formats show good performance in LLM quantization, they tend to perform poorly with small group sizes or sub-4 bits. We find the reason is that the absence

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).