← all papers · overview

Fast-slow Thinking RM: Efficient Integration Of Scalar And Generative Reward Models

Abstract

Reward models (RMs) are critical for aligning Large Language Models via Reinforcement Learning from Human Feedback (RLHF). While Generative Reward Models (GRMs) achieve superior accuracy through chain-of-thought (CoT) reasoning, they incur substantial computational costs. Conversely, Scalar Reward Models (SRMs) offer efficiency but suffer from limited performance and adaptability in complex scenar

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).