← all papers · overview

Toward Inference-optimal Mixture-of-expert Large Language Models

Abstract

Mixture-of-Expert (MoE) based large language models (LLMs), such as the recent Mixtral and DeepSeek-MoE, have shown great promise in scaling model size without suffering from the quadratic growth of training cost of dense transformers. Like dense models, training MoEs requires answering the same question: given a training budget, what is the optimal allocation on the model size and number of token

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).