← all papers · overview

Konkani LLM: Multi-script Instruction Tuning And Evaluation For A Low-resource Indian Language

Abstract

Large Language Models (LLMs) consistently under perform in low-resource linguistic contexts such as Konkani. This performance deficit stems from acute training data scarcity compounded by high script diversity across Devanagari, Romi and Kannada orthographies. To address this gap, we introduce Konkani-Instruct-100k, a comprehensive synthetic instruction-tuning dataset generated through Gemini 3.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).