← all papers · overview

Alloc-moe: Budget-aware Expert Activation Allocation For Efficient Mixture-of-experts Inference

Abstract

Mixture-of-Experts (MoE) has become a dominant architecture for scaling large language models due to their sparse activation mechanism. However, the substantial number of expert activations creates a critical latency bottleneck during inference, especially in resource-constrained deployment scenarios. Existing approaches that reduce expert activations potentially lead to severe model performance d

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).