← all papers · overview

An Imperfect Verifier Is Good Enough: Learning With Noisy Rewards

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has become a prominent method for post-training Large Language Models (LLMs). However, verifiers are rarely error-free; even deterministic checks can be inaccurate, and the growing dependence on model-based judges exacerbates the issue. The extent to which RLVR is robust to such noise and the verifier accuracy required for effective training re

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).