← all papers · overview

MINI-LLM: Memory-efficient Structured Pruning For Large Language Models

Abstract

As Large Language Models (LLMs) grow dramatically in size, there is an increasing trend in compressing and speeding up these models. Previous studies have highlighted the usefulness of gradients for importance scoring in neural network compressing, especially in pruning medium-size networks. However, the substantial memory requirements involved in calculating gradients with backpropagation impede

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).