← all papers · overview

Romansetu: Efficiently Unlocking Multilingual Capabilities Of Large Language Models Via Romanization

Abstract

This study addresses the challenge of extending Large Language Models (LLMs) to non-English languages that use non-Roman scripts. We propose an approach that utilizes the romanized form of text as an interface for LLMs, hypothesizing that its frequent informal use and shared tokens with English enhance cross-lingual alignment. Our approach involves the continual pretraining of an English LLM like

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).