← all papers · overview

On The Algorithmic Bias Of Aligning Large Language Models With RLHF: Preference Collapse And Matching Regularization

Abstract

Accurately aligning large language models (LLMs) with human preferences is crucial for informing fair, economically sound, and statistically efficient decision-making processes. However, we argue that the predominant approach for aligning LLMs with human preferences through a reward model -- reinforcement learning from human feedback (RLHF) -- suffers from an inherent algorithmic bias due to its K

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).