← all papers · overview

Neuronmoe: Neuron-guided Mixture-of-experts For Efficient Multilingual LLM Extension

Abstract

Extending large language models to low-resource languages is essential for global accessibility, but training separate models per language is prohibitively expensive. Mixture-of-Experts (MoE) architectures address this by adding sparse language-specific parameters, but determining how many experts each layer needs remains an open question. Current approaches allocate experts based on layer-level s

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).