← all papers · overview

Accelerating LLM Pre-training Through Flat-direction Dynamics Enhancement

Abstract

Pre-training Large Language Models requires immense computational resources, making optimizer efficiency essential. The optimization landscape is highly anisotropic, with loss reduction driven predominantly by progress along flat directions. While matrix-based optimizers such as Muon and SOAP leverage fine-grained curvature information to outperform AdamW, their updates tend toward isotropy -- rel

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).