← all papers · overview

QUAD: Quantization And Parameter-efficient Tuning Of LLM With Activation Decomposition

Abstract

Large Language Models (LLMs) excel in diverse applications but suffer inefficiency due to massive scale. While quantization reduces computational costs, existing methods degrade accuracy in medium-sized LLMs (e.g., Llama-3-8B) due to activation outliers. To address this, we propose QUAD (Quantization with Activation Decomposition), a framework leveraging Singular Value Decomposition (SVD) to suppr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).