← all papers · overview

Moe-lpr: Multilingual Extension Of Large Language Models Through Mixture-of-experts With Language Priors Routing

Abstract

Large Language Models (LLMs) are often English-centric due to the disproportionate distribution of languages in their pre-training data. Enhancing non-English language capabilities through post-pretraining often results in catastrophic forgetting of the ability of original languages. Previous methods either achieve good expansion with severe forgetting or slight forgetting with poor expansion, ind

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).