← all papers · overview

Weight Decay Improves Language Model Plasticity

Abstract

The prevailing paradigm in large language model (LLM) development is to pretrain a base model, then perform further training to improve performance and model behavior. However, hyperparameter optimization and scaling laws have been studied primarily from the perspective of the base model's validation loss, ignoring downstream adaptability. In this work, we study pretraining from the perspective of

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).