← all papers · overview

Uncertainty-penalized Reinforcement Learning From Human Feedback With Diverse Reward Lora Ensembles

Abstract

Reinforcement learning from human feedback (RLHF) emerges as a promising paradigm for aligning large language models (LLMs). However, a notable challenge in RLHF is overoptimization, where beyond a certain threshold, the pursuit of higher rewards leads to a decline in human preferences. In this paper, we observe the weakness of KL regularization which is commonly employed in existing RLHF methods

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).