← all papers · overview

Chinese Tiny LLM: Pretraining A Chinese-centric Large Language Model

Abstract

In this study, we introduce CT-LLM, a 2B large language model (LLM) that illustrates a pivotal shift towards prioritizing the Chinese language in developing LLMs. Uniquely initiated from scratch, CT-LLM diverges from the conventional methodology by primarily incorporating Chinese textual data, utilizing an extensive corpus of 1,200 billion tokens, including 800 billion Chinese tokens, 300 billion

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).