← all papers · overview

Spannorm: Reconciling Training Stability And Performance In Deep Transformers

Abstract

The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures. A critical design choice is the placement of normalization layers, leading to a fundamental trade-off: the ``PreNorm'' architecture ensures training stability at the cost of potential performance degradation in deep models, while the ``PostNorm'' architecture offers strong performance but s

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).