← all papers · overview

Mano: Restriking Manifold Optimization For LLM Training

Abstract

While large language models (LLMs) have emerged as a significant advancement in artificial intelligence, the hardware and computational costs for training LLMs are also significantly burdensome. Among the state-of-the-art optimizers, AdamW relies on diagonal curvature estimates and ignores structural properties, while Muon applies global spectral normalization at the expense of losing curvature in

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).