← all papers · overview

HLAT: High-quality Large Language Model Pre-trained On AWS Trainium

Abstract

Getting large language models (LLMs) to perform well on the downstream tasks requires pre-training over trillions of tokens. This typically demands a large number of powerful computational devices in addition to a stable distributed training framework to accelerate the training. The growing number of applications leveraging AI/ML led to a scarcity of the expensive conventional accelerators (such a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).