← all papers · overview

Locmoe: A Low-overhead Moe For Large Language Model Training

Abstract

The Mixtures-of-Experts (MoE) model is a widespread distributed and integrated learning method for large language models (LLM), which is favored due to its ability to sparsify and expand models efficiently. However, the performance of MoE is limited by load imbalance and high latency of All-to-All communication, along with relatively redundant computation owing to large expert capacity. Load imbal

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).