Abstract
Large Language Models (LLMs) pose significant hardware challenges related to memory requirements and computational ability. There are two mainstream quantization schemes for LLMs: coarse-grained ( channel-wise) quantization and fine-grained ( group-wise) quantization. Fine-grained quantization has smaller quantization loss, consequently achieving superior pe