← all papers · overview

Mixed Sparsity Training: Achieving 4 FLOP Reduction For Transformer Pretraining

Abstract

Large language models (LLMs) have made significant strides in complex tasks, yet their widespread adoption is impeded by substantial computational demands. With hundreds of billion parameters, transformer-based LLMs necessitate months of pretraining across a high-end GPU cluster. However, this paper reveals a compelling finding: transformers exhibit considerable redundancy in pretraining computati

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).