← all papers · overview

Reward Models Inherit Value Biases From Pretraining

Abstract

Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pretrained and post-trained LLMs themselves. Because RMs are initialized from LLMs, they inherit representations that shape their behavior, but the nature and extent of this influence remain understudied. In a comprehensive study of 10 leading open-weight RMs using valid

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).