← all papers · overview

GPTQT: Quantize Large Language Models Twice To Push The Efficiency

Abstract

Due to their large size, generative Large Language Models (LLMs) require significant computing and storage resources. This paper introduces a new post-training quantization method, GPTQT, to reduce memory usage and enhance processing speed by expressing the weight of LLM in 3bit/2bit. Practice has shown that minimizing the quantization error of weights is ineffective, leading to overfitting. There

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).