← all papers · overview

TFD: A Comprehensive Structured Tibetan Foundation Dataset For Low-resource Language Processing And Large-scale Modeling

Abstract

Large language models (LLMs) have achieved remarkable success in high-resource languages, yet progress for Tibetan remains severely constrained by the lack of large-scale, high-quality, and structured data. Existing Tibetan resources are fragmented, domain-limited, and insufficient to support modern LLM pipelines requiring pretraining, instruction tuning, safety alignment, and reasoning supervisio

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).