← all papers · overview

Bayesian Reward Models For LLM Alignment

Abstract

To ensure that large language model (LLM) responses are helpful and non-toxic, a reward model trained on human preference data is usually used. LLM responses with high rewards are then selected through best-of- (BoN) sampling or the LLM is further optimized to produce responses with high rewards through reinforcement learning from human feedback (RLHF). However, these processes are susceptibl

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).