← all papers · overview

Moe Routing Testbed: Studying Expert Specialization And Routing Behavior At Small Scale

Abstract

Sparse Mixture-of-Experts (MoE) architectures are increasingly popular for frontier large language models (LLM) but they introduce training challenges due to routing complexity. Fully leveraging parameters of an MoE model requires all experts to be well-trained and to specialize in non-redundant ways. Assessing this, however, is complicated due to lack of established metrics and, importantly, many

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).