← all papers · overview

The Fine-grained Complexity Of Gradient Computation For Training Large Language Models

Abstract

Large language models (LLMs) have made fundamental contributions over the last a few years. To train an LLM, one needs to alternatingly run `forward' computations and `backward' computations. The forward computation can be viewed as attention function evaluation, and the backward computation can be viewed as a gradient computation. In previous work by [Alman and Song, NeurIPS 2023], it was proved

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).