← all papers · overview

Aboutme: Using Self-descriptions In Webpages To Document The Effects Of English Pretraining Data Filters

Abstract

Large language models' (LLMs) abilities are drawn from their pretraining data, and model development begins with data curation. However, decisions around what data is retained or removed during this initial stage are under-scrutinized. In our work, we ground web text, which is a popular pretraining data source, to its social and geographic contexts. We create a new dataset of 10.3 million self-des

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).