← all papers · overview

Pre-training LLM Without Learning Rate Decay Enhances Supervised Fine-tuning

Abstract

We investigate the role of learning rate scheduling in the large-scale pre-training of large language models, focusing on its influence on downstream performance after supervised fine-tuning (SFT). Decay-based learning rate schedulers are widely used to minimize pre-training loss. However, despite their widespread use, how these schedulers affect performance after SFT remains underexplored. In thi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).