← all papers · overview

Efficient Post-training Pruning Of Large Language Models With Statistical Correction

Abstract

Post-training pruning is an effective approach for reducing the size and inference cost of large language models (LLMs), but existing methods often face a trade-off between pruning quality and computational efficiency. Heuristic pruning methods are efficient but sensitive to activation outliers, while reconstruction-based approaches improve fidelity at the cost of heavy computation. In this work,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).