← all papers · overview

Progressive Residual Warmup For Language Model Pretraining

Abstract

Transformer architectures serve as the backbone for most modern Large Language Models, therefore their pretraining stability and convergence speed are of central concern. Motivated by the logical dependency of sequentially stacked layers, we propose Progressive Residual Warmup (ProRes) for language model pretraining. ProRes implements an "early layer learns first" philosophy by multiplying each la

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).