← all papers · overview

How Long Reasoning Chains Influence Llms' Judgment Of Answer Factuality

Abstract

Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible to surface-level biases. One possible reason is that these judges lack sufficient information in assessing answer correctness. With the rise of reasoning-capable models, exposing a generator's reasoning content to the judge provides richer information

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).