← all papers · overview

Funprm: Function-as-step Process Reward Model With Meta Reward Correction For Code Generation

Abstract

Code generation is a core application of large language models (LLMs), yet LLMs still frequently fail on complex programming tasks. Given its success in mathematical reasoning, test-time scaling approaches such as Process Reward Model (PRM)-based Best-of-N selection offer a promising way to improve performance. However, existing PRMs remain ineffective for code generation due to the lack of meanin

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).