← all papers · overview

On The Global Convergence Of Online RLHF With Neural Parametrization

Abstract

The importance of Reinforcement Learning from Human Feedback (RLHF) in aligning large language models (LLMs) with human values cannot be overstated. RLHF is a three-stage process that includes supervised fine-tuning (SFT), reward learning, and policy learning. Although there are several offline and online approaches to aligning LLMs, they often suffer from distribution shift issues. These issues a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).