← all papers · overview

Regularized Best-of-n Sampling With Minimum Bayes Risk Objective For Language Model Alignment

Abstract

Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) to human preferences at the time of decoding. BoN sampling is susceptible to a problem known as reward hacking when the accuracy of the reward model is not high enough due to the quality or the quantity of the preference dataset. Because the reward model is an imperfect

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).