← all papers · overview

Scalebits: Scalable Bitwidth Search For Hardware-aligned Mixed-precision Llms

Abstract

Post-training weight quantization is crucial for reducing the memory and inference cost of large language models (LLMs), yet pushing the average precision below 4 bits remains challenging due to highly non-uniform weight sensitivity and the lack of principled precision allocation. Existing solutions use irregular fine-grained mixed-precision with high runtime overhead or rely on heuristics or high

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).