← all papers · overview

Prompt-prompted Adaptive Structured Pruning For Efficient LLM Generation

Abstract

With the development of transformer-based large language models (LLMs), they have been applied to many fields due to their remarkable utility, but this comes at a considerable computational cost at deployment. Fortunately, some methods such as pruning or constructing a mixture of experts (MoE) aim at exploiting sparsity in transformer feedforward (FF) blocks to gain boosts in speed and reduction i

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).