← all papers · overview

The Alignment Ceiling: Objective Mismatch In Reinforcement Learning From Human Feedback

Abstract

Reinforcement learning from human feedback (RLHF) has emerged as a powerful technique to make large language models (LLMs) more capable in complex settings. RLHF proceeds as collecting human preference data, training a reward model on said data, and optimizing a base ML model with respect to said reward for extrinsic evaluation metrics (e.g. MMLU, GSM8k). RLHF relies on many assumptions about how

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).