← all papers · overview

Pruning Foundation Models For High Accuracy Without Retraining

Abstract

Despite the superior performance, it is challenging to deploy foundation models or large language models (LLMs) due to their massive parameters and computations. While pruning is a promising technique to reduce model size and accelerate the inference, the traditional pruning techniques can hardly be applied for LLMs as they need to finetune the model on the full dataset with multiple epochs consum

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).