← all papers · overview

You Only Prune Once: Designing Calibration-free Model Compression With Policy Learning

Abstract

The ever-increasing size of large language models (LLMs) presents significant challenges for deployment due to their heavy computational and memory requirements. Current model pruning techniques attempt to alleviate these issues by relying heavily on external calibration datasets to determine which parameters to prune or compress, thus limiting their flexibility and scalability across different co

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).