← all papers · overview

What Drives Performance In Multilingual Language Models?

Abstract

This study investigates the factors influencing the performance of multilingual large language models (MLLMs) across diverse languages. We study 6 MLLMs, including masked language models, autoregressive models, and instruction-tuned LLMs, on the SIB-200 dataset, a topic classification dataset encompassing 204 languages. Our analysis considers three scenarios: ALL languages, SEEN languages (present

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).