← all papers · overview

BPDQ: Bit-plane Decomposition Quantization On A Variable Grid For Large Language Models

Abstract

Large language model (LLM) inference is often bounded by memory footprint and memory bandwidth in resource-constrained deployments, making quantization a fundamental technique for efficient serving. While post-training quantization (PTQ) maintains high fidelity at 4-bit, it deteriorates at 2-3 bits. Fundamentally, existing methods enforce a shape-invariant quantization grid (e.g., the fixed unifor

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).