← all papers · overview

TEON: Tensorized Orthonormalization Beyond Layer-wise Muon For Large Language Model Pre-training

Abstract

The Muon optimizer has demonstrated strong empirical performance in pre-training large language models by performing matrix-level gradient (or momentum) orthogonalization in each layer independently. In this work, we propose TEON, a principled generalization of Muon that extends orthogonalization beyond individual layers by modeling the gradients of a neural network as a structured higher-order te

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).