← all papers · overview

QQQ: Quality Quattuor-bit Quantization For Large Language Models

Abstract

Quantization is a proven effective method for compressing large language models. Although popular techniques like W8A8 and W4A16 effectively maintain model performance, they often fail to concurrently speed up the prefill and decoding stages of inference. W4A8 is a promising strategy to accelerate both of them while usually leads to a significant performance degradation. To address these issues, w

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).