← all papers · overview

DB-LLM: Accurate Dual-binarization For Efficient Llms

Abstract

Large language models (LLMs) have significantly advanced the field of natural language processing, while the expensive memory and computation consumption impede their practical deployment. Quantization emerges as one of the most effective methods for improving the computational efficiency of LLMs. However, existing ultra-low-bit quantization always causes severe accuracy drops. In this paper, we e

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).