← all papers · overview

Beyond Binary Preferences: A Principled Framework For Reward Modeling With Ordinal Feedback

Abstract

Reward modeling is crucial for aligning large language models with human preferences, yet current approaches lack a principled mathematical framework for leveraging ordinal preference data. When human annotators provide graded preferences on a Likert scale (e.g., significantly better, better, slightly better, negligibly better), existing methods typically apply ad-hoc heuristics, such as margin te

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).