← all papers · overview

Stochastic Rounding For LLM Training: Theory And Practice

Abstract

As the parameters of Large Language Models (LLMs) have scaled to hundreds of billions, the demand for efficient training methods -- balancing faster computation and reduced memory usage without sacrificing accuracy -- has become more critical than ever. In recent years, various mixed precision strategies, which involve different precision levels for optimization components, have been proposed to i

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).