← all papers · overview

Beyond Model Collapse: Scaling Up With Synthesized Data Requires Verification

Abstract

Large Language Models (LLM) are increasingly trained on data generated by other LLM, either because generated text and images become part of the pre-training corpus, or because synthetized data is used as a replacement for expensive human-annotation. This raises concerns about *model collapse*, a drop in model performance when their training sets include generated data. Considering that it is easi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).