← all papers · overview

Muon+: Towards Better Muon Via One Additional Normalization Step

Abstract

The Muon optimizer has demonstrated promising performance in pre-training large language models through gradient (or momentum) orthogonalization. In this work, we propose a simple yet effective enhancement to Muon, namely Muon+, which introduces an additional normalization step after orthogonalization. We demonstrate the effectiveness of Muon+ through extensive pre-training experiments across a wi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).