← all papers · overview

A Unified Pairwise Framework For RLHF: Bridging Generative Reward Modeling And Policy Optimization

Abstract

Reinforcement Learning from Human Feedback (RLHF) has emerged as a important paradigm for aligning large language models (LLMs) with human preferences during post-training. This framework typically involves two stages: first, training a reward model on human preference data, followed by optimizing the language model using reinforcement learning algorithms. However, current RLHF approaches may cons

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).