← all papers · overview

STUN: Structured-then-unstructured Pruning For Scalable Moe Pruning

Abstract

Mixture-of-experts (MoEs) have been adopted for reducing inference costs by sparsely activating experts in Large language models (LLMs). Despite this reduction, the massive number of experts in MoEs still makes them expensive to serve. In this paper, we study how to address this, by pruning MoEs. Among pruning methodologies, unstructured pruning has been known to achieve the highest performance fo

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).