← all papers · overview

A Baseline Analysis Of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift

Abstract

Foundation models, specifically Large Language Models (LLMs), have lately gained wide-spread attention and adoption. Reinforcement Learning with Human Feedback (RLHF) involves training a reward model to capture desired behaviors, which is then used to align LLM's. These reward models are additionally used at inference-time to estimate LLM responses' adherence to those desired behaviors. However, t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).