← all papers · overview

Does A Global Perspective Help Prune Sparse Moes Elegantly?

Abstract

Empirical scaling laws for language models have encouraged the development of ever-larger LLMs, despite their growing computational and memory costs. Sparse Mixture-of-Experts (MoEs) offer a promising alternative by activating only a subset of experts per forward pass, improving efficiency without sacrificing performance. However, the large number of expert parameters still leads to substantial me

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).