← all papers · overview

Gneissweb: Preparing High Quality Data For Llms At Scale

Abstract

Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of downstream tasks. Large pre-training datasets for leading LLMs remain inaccessible to the public, whereas many open datasets are small in size (less than 5 trillion tokens), limiting

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).