← all papers · overview

Optimal Low-rank Stochastic Gradient Estimation For LLM Training

Abstract

Large language model (LLM) training is often bottlenecked by memory constraints and stochastic gradient noise in extremely high-dimensional parameter spaces. Motivated by empirical evidence that many LLM gradient matrices are effectively low-rank during training, we present an unbiased, memory-efficient, low-rank matrix estimator with the lowest variance that is applicable across common stochastic

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).