← all papers · overview

Lorap: Transformer Sub-layers Deserve Differentiated Structured Compression For Large Language Models

Abstract

Large language models (LLMs) show excellent performance in difficult tasks, but they often require massive memories and computational resources. How to reduce the parameter scale of LLMs has become research hotspots. In this study, we make an important observation that the multi-head self-attention (MHA) sub-layer of Transformer exhibits noticeable low-rank structure, while the feed-forward networ

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).