← all papers · overview

Scaling Law For Language Models Training Considering Batch Size

Abstract

Large language models (LLMs) have made remarkable advances in recent years, with scaling laws playing a critical role in this rapid progress. In this paper, we empirically investigate how a critical hyper-parameter, i.e., the global batch size, influences the LLM training prdocess. We begin by training language models ranging from 125 million to 2.6 billion parameters, using up to 300 billion high

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).