← all papers · overview

Gprune-llm: Generalization-aware Structured Pruning For Large Language Models

Abstract

Structured pruning is widely used to compress large language models (LLMs), yet its effectiveness depends heavily on neuron importance estimation. Most existing methods estimate neuron importance from activation statistics on a single calibration dataset, which introduces calibration bias and degrades downstream cross-task generalization. We observe that neurons exhibit heterogeneous distribution

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).