← all papers · overview

Robust Reward Modeling For Large Language Models Via Causal Decomposition

Abstract

Reward models are central to aligning large language models, yet they often overfit to spurious cues such as response length and overly agreeable tone. Most prior work weakens these cues directly by penalizing or controlling specific artifacts, but it does not explicitly encourage the model to ground preferences in the prompt's intent. We learn a de

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).