← all papers · overview

Protrain: Efficient LLM Training Via Memory-aware Techniques

Abstract

It is extremely memory-hungry to train Large Language Models (LLM). To solve this problem, existing work exploits the combination of CPU and GPU for the training process, such as ZeRO-Offload. Such a technique largely democratizes billion-scale model training, making it possible to train with few consumer graphics cards. However, based on our observation, existing frameworks often provide coarse-g

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).