← all papers · overview

Zyda: A 1.3T Dataset For Open Language Modeling

Abstract

The size of large language models (LLMs) has scaled dramatically in recent years and their computational and data requirements have surged correspondingly. State-of-the-art language models, even at relatively smaller sizes, typically require training on at least a trillion tokens. This rapid advancement has eclipsed the growth of open-source datasets available for large-scale LLM pretraining. In t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).