← all papers · overview

SPQ: An Ensemble Technique For Large Language Model Compression

Abstract

This study presents an ensemble technique, SPQ (SVD-Pruning-Quantization), for large language model (LLM) compression that combines variance-retained singular value decomposition (SVD), activation-based pruning, and post-training linear quantization. Each component targets a different source of inefficiency: i) pruning removes redundant neurons in MLP layers, ii) SVD reduces attention projections

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).