← all papers · overview

To FP8 And Back Again: Quantifying Reduced Precision Effects On LLM Training Stability

Abstract

The massive computational costs associated with large language model (LLM) pretraining have spurred great interest in reduced-precision floating-point representations to accelerate the process. As a result, the BrainFloat16 (BF16) precision has become the de facto standard for LLM training, with hardware support included in recent generations of accelerators. This trend has gone even further in th

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).