← all papers · overview

An Empirical Study On Noisy Data And LLM Pretraining Loss Divergence

Abstract

Large-scale pretraining datasets drive the success of large language models (LLMs). However, these web-scale corpora inevitably contain large amounts of noisy data due to unregulated web content or randomness inherent in data. Although LLM pretrainers often speculate that such noise contributes to instabilities in large-scale LLM pretraining and, in the worst cases, loss divergence, this phenomeno

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).