← all papers · overview

HESTIA: A Hessian-guided Differentiable Quantization-aware Training Framework For Extremely Low-bit Llms

Abstract

As large language models (LLMs) continue to scale, deployment is increasingly bottlenecked by the memory wall, motivating a shift toward extremely low-bit quantization. However, most quantization-aware training (QAT) methods apply hard rounding and the straight-through estimator (STE) from the beginning of the training, which prematurely discretizes the optimization landscape and induces persisten

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).