← all papers · overview

An Adaptive Placement And Parallelism Framework For Accelerating RLHF Training

Abstract

Recently, ChatGPT or InstructGPT like large language models (LLM) has made a significant impact in the AI world. Many works have attempted to reproduce the complex InstructGPT's training pipeline, namely Reinforcement Learning with Human Feedback (RLHF). However, the mainstream distributed RLHF training methods typically adopt a fixed model placement strategy, referred to as the Co-located strateg

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).