← all papers · overview

Beyond The Last Answer: Your Reasoning Trace Uncovers More Than You Think

Abstract

Large Language Models (LLMs) leverage step-by-step reasoning to solve complex problems. Standard evaluation practice involves generating a complete reasoning trace and assessing the correctness of the final answer presented at its conclusion. In this paper, we challenge the reliance on the final answer by posing the following two questions: Does the final answer reliably represent the model's opti

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).