← all papers · overview

Regurgitative Training: The Value Of Real Data In Training Large Language Models

Abstract

What happens if we train a new Large Language Model (LLM) using data that are at least partially generated by other LLMs? The explosive success of LLMs means that a substantial amount of content online will be generated by LLMs rather than humans, which will inevitably enter the training datasets of next-generation LLMs. We evaluate the implications of such "regurgitative training" on LLM performa

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).