← all papers · overview

Building Pre-train LLM Dataset For The INDIC Languages: A Case Study On Hindi

Abstract

Large language models (LLMs) demonstrated transformative capabilities in many applications that require automatically generating responses based on human instruction. However, the major challenge for building LLMs, particularly in Indic languages, is the availability of high-quality data for building foundation LLMs. In this paper, we are proposing a large pre-train dataset in Hindi useful for the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).