← all papers · overview

Indicllmsuite: A Blueprint For Creating Pre-training And Fine-tuning Datasets For Indian Languages

Abstract

Despite the considerable advancements in English LLMs, the progress in building comparable models for other languages has been hindered due to the scarcity of tailored resources. Our work aims to bridge this divide by introducing an expansive suite of resources specifically designed for the development of Indic LLMs, covering 22 languages, containing a total of 251B tokens and 74.8M instruction-re

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).