← all papers · overview

Beware Of Calibration Data For Pruning Large Language Models

Abstract

As large language models (LLMs) are widely applied across various fields, model compression has become increasingly crucial for reducing costs and improving inference efficiency. Post-training pruning is a promising method that does not require resource-intensive iterative training and only needs a small amount of calibration data to assess the importance of parameters. Recent research has enhance

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).