← all papers · overview

Optimizing Distributed Training On Frontier For Large Language Models

Abstract

Large language models (LLMs) have demonstrated remarkable success as foundational models, benefiting various downstream applications through fine-tuning. Recent studies on loss scaling have demonstrated the superior performance of larger LLMs compared to their smaller counterparts. Nevertheless, training LLMs with billions of parameters poses significant challenges and requires considerable comput

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).