← all papers · overview

Process Supervision For Chain-of-thought Reasoning Via Monte Carlo Net Information Gain

Abstract

Multi-step reasoning improves the capabilities of large language models (LLMs) but increases the risk of errors propagating through intermediate steps. Process reward models (PRMs) mitigate this by scoring each step individually, enabling fine-grained supervision and improved reliability. Existing methods for training PRMs rely on costly human annotations or computationally intensive automatic lab

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).