← all papers · overview

Cdquant: Greedy Coordinate Descent For Accurate LLM Quantization

Abstract

Large language models (LLMs) have recently demonstrated remarkable performance across diverse language tasks. But their deployment is often constrained by their substantial computational and storage requirements. Quantization has emerged as a key technique for addressing this challenge, enabling the compression of large models with minimal impact on performance. The recent GPTQ algorithm, a post-t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).