← all papers · overview

Elephant In The Room: Unveiling The Impact Of Reward Model Quality In Alignment

Abstract

The demand for regulating potentially risky behaviors of large language models (LLMs) has ignited research on alignment methods. Since LLM alignment heavily relies on reward models for optimization or evaluation, neglecting the quality of reward models may cause unreliable results or even misalignment. Despite the vital role reward models play in alignment, previous works have consistently overloo

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).