← all papers · overview

Generating Data-driven Reasoning Rubrics For Domain-adaptive Reward Modeling

Abstract

An impediment to using Large Language Models (LLMs) for reasoning output verification is that LLMs struggle to reliably identify errors in thinking traces, particularly in long outputs, domains requiring expert knowledge, and problems without verifiable rewards. We propose a data-driven approach to automatically construct highly granular reasoning error taxonomies to enhance LLM-driven error detec

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).