← all papers · overview

: A Generalist Value Model For Any Policy At State Zero

Abstract

Policy gradient methods rely on a baseline to measure the relative advantage of an action, ensuring the model reinforces behaviors that outperform its current average capability. In the training of Large Language Models (LLMs) using Actor-Critic methods (e.g., PPO), this baseline is typically estimated by a Value Model (Critic) often as large as the policy model itself. However, as the policy cont

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).