← all papers · overview

When Less Is More: Investigating Data Pruning For Pretraining Llms At Scale

Abstract

Large volumes of text data have contributed significantly to the development of large language models (LLMs) in recent years. This data is typically acquired by scraping the internet, leading to pretraining datasets comprised of noisy web text. To date, efforts to prune these datasets down to a higher quality subset have relied on hand-crafted heuristics encoded as rule-based filters. In this work

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).