← all papers · overview

Language-specific Neurons: The Key To Multilingual Capabilities In Large Language Models

Abstract

Large language models (LLMs) demonstrate remarkable multilingual capabilities without being pre-trained on specially curated multilingual parallel corpora. It remains a challenging problem to explain the underlying mechanisms by which LLMs process multilingual texts. In this paper, we delve into the composition of Transformer architectures in LLMs to pinpoint language-specific regions. Specially,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).