← all papers · overview

Scalable Pretraining Of Large Mixture Of Experts Language Models On Aurora Super Computer

Abstract

Pretraining Large Language Models (LLMs) from scratch requires massive amount of compute. Aurora super computer is an ExaScale machine with 127,488 Intel PVC (Ponte Vechio) GPU tiles. In this work, we showcase LLM pretraining on Aurora at the scale of 1000s of GPU tiles. Towards this effort, we developed Optimus, an inhouse training library with support for standard large model training techniques

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).