← all papers · overview

From Calculation To Adjudication: Examining LLM Judges On Mathematical Reasoning Tasks

Abstract

To reduce the need for human annotations, large language models (LLMs) have been proposed as judges of the quality of other candidate models. The performance of LLM judges is typically evaluated by measuring the correlation with human judgments on generative tasks such as summarization or machine translation. In contrast, we study LLM judges on mathematical reasoning tasks. These tasks require mul

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).