← all papers · overview

TRANSOM: An Efficient Fault-tolerant System For Training Llms

Abstract

Large language models (LLMs) with hundreds of billions or trillions of parameters, represented by chatGPT, have achieved profound impact on various fields. However, training LLMs with super-large-scale parameters requires large high-performance GPU clusters and long training periods lasting for months. Due to the inevitable hardware and software failures in large-scale clusters, maintaining uninte

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).