← all papers · overview

Deterministic Differentiable Structured Pruning For Large Language Models

Abstract

Structured pruning reduces LLM inference cost by removing low-importance architectural components. This can be viewed as learning a multiplicative gate for each component under an l0 sparsity constraint. Due to the discreteness of the l0 norm, prior work typically adopts stochastic hard-concrete relaxations to enable differentiable optimization; however, this stochasticity can introduce a train--t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).