← all papers · overview

Tbdfiltering: Sample-efficient Tree-based Data Filtering

Abstract

The quality of machine learning models depends heavily on their training data. Selecting high-quality, diverse training sets for large language models (LLMs) is a difficult task, due to the lack of cheap and reliable quality metrics. While querying existing LLMs for document quality is common, this is not scalable to the large number (billions) of documents used in training. Instead, practitioners

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).