← all papers · overview

Mixture Compressor For Mixture-of-experts Llms Gains More

Abstract

Mixture-of-Experts large language models (MoE-LLMs) marks a significant step forward of language models, however, they encounter two critical challenges in practice: 1) expert parameters lead to considerable memory consumption and loading latency; and 2) the current activated experts are redundant, as many tokens may only require a single expert. Motivated by these issues, we investigate the MoE-L

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).